Experiment / E621JF7G9Standard STARR-seq

PANC-1 STARR-seq screen of PDAC non-coding mutations

Non-coding mutations at enhancer clusters contribute to pancreatic ductal adenocarcinoma

An episomal hSTARR-seq library tested 587 candidate PDAC somatic non-coding mutations and corresponding wild-type sequences in PANC-1 cells using 230-bp oligos with five shifted genomic windows per allele, plus positive and negative controls. Two biological replicates were analyzed; the processed table contains the 43 unique mutations represented by the 48 top-significant oligos listed in Extended Data S5 (p<0.01).

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Synthetic candidate oligos were cloned into the second-generation hSTARR-seq ORI plasmid (Addgene #99296), downstream of the core promoter in the reporter 3′ UTR. Each construct contained a unique 6-bp barcode and a 194-bp genomic sequence context flanked by 15-bp linker regions, for a nominal 230-bp construct. Mutant and wild-type alleles were each represented in five contexts: left_20, left_10, center, right_10, and right_20. PANC-1 cells were transfected with the plasmid pool, poly(A)+ reporter RNA and plasmid DNA were sequenced, and activity was calculated as normalized RNA TPM divided by DNA TPM.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (19 of 19)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 19 definitions
variant_id
Normalized hg19 variant identifier in the form chromosome:position:reference>alternate; '-' denotes an insertion/deletion allele in the source notation.
chromosome
Human chromosome in hg19 notation.
position_hg19
1-based hg19 coordinate of the mutation as reported in the oligo name.
ref_allele
Reference allele or deleted sequence from the source oligo name.
alt_allele
Alternate allele or inserted sequence from the source oligo name.
variant_class
SNV, insertion, deletion, or other indel classification based on the reported alleles.
annotated_gene_or_element
Gene or regulatory element label embedded in the Extended Data S5 oligo name.
top_hit_oligo_ids
One or more source oligo labels (M1–M48) representing this mutation among the reported top hits; pipe-delimited when multiple windows are listed.
top_hit_windows
Source sliding-window labels for the top-hit oligos: L_20, L, M, R, or R_20.
top_hit_window_offsets_bp
Nominal mutation-window displacement derived from the source label: -20, -10, 0, 10, or 20 bp.
top_hit_oligo_count
Number of top-significant oligos for this mutation listed in Extended Data S5.
reporter_effect_direction
Mutant reporter activity direction relative to wild type, mapped from the gain/loss bar colors in Fig. 2e.
reported_significance
Significance threshold stated in the Extended Data S5 caption; exact per-oligo p-values were not supplied in the public workbook.
predicted_tf_motif
Transcription-factor motif explicitly discussed for this mutation in the paper text/figure; blank when not explicitly reported.
motif_change
Explicitly reported motif gain or motif break for the listed transcription factor; blank when not reported.
reported_variant_mean_log2_fc
Exact variant-level mean log2 fold change reported in the paper text for the small number of highlighted mutations; blank when not publicly reported.
reported_variant_p_value
Exact or inequality p-value for the corresponding text-reported variant-level summary; blank when not publicly reported.
reported_effect_summary_note
Scope of any text-reported effect summary, including the number of sliding windows and test description.
qc_status
PASS indicates inclusion in the authors’ top-significant-oligo list after the stated sequencing/UMI QC scope.

Quality control

The authors report low-quality read filtering, adapter/linker trimming, 6-bp barcode-based separation of WT, MUT, and control constructs, alignment to the hg19 oligo design, UMI deduplication, and a minimum of 3 unique UMIs per counted construct. The cloned plasmid library had minimum 30× coverage for 98.63% of the 6,082 sequenced constructs with WT and MUT represented. Two biological replicates were retained for the STARR-seq screen. The processed table is further restricted to the authors’ Extended Data S5 list of top-significant oligos (p<0.01), so all retained rows pass the reported screen-level significance/QC scope.

Curation notes

The complete STARR-seq RNA/DNA count matrix, per-oligo activity values, and exact p-values were not deposited publicly; the authors state that STARR-seq data and analysis scripts can be requested from the corresponding authors. Accordingly, table.csv is a transparent derived summary of the public Extended Data S5 top-hit list, not a complete 6,082-oligo assay matrix. The 48 source oligos were grouped into 43 unique mutations, matching the abstract’s 43 prioritized NCMs. Reporter direction was read from Fig. 2e (M1–M37 gain of function; M38–M48 loss of function), and the resulting unique-variant counts are 33 gain and 10 loss. Exact text-reported summaries are retained for chr12:115067012 C>A (mean log2 fold change 3.69 across five sliding windows, p=0.016) and chr3:71123616 G>T (average log2 fold change -1.36 across three significant windows, p<0.05). The paper alternates between 6,080 constructs and 6,082 oligos/constructs in different sections; the table uses the 6,082 value from the library description and QC statement. The source workbook labels the chr12:115067012 C>A region TBX5-AS1, whereas the text describes it as proximal to TBX3; the source label is preserved in the table.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.