PANC-1 STARR-seq screen of PDAC non-coding mutations
Non-coding mutations at enhancer clusters contribute to pancreatic ductal adenocarcinomaAn episomal hSTARR-seq library tested 587 candidate PDAC somatic non-coding mutations and corresponding wild-type sequences in PANC-1 cells using 230-bp oligos with five shifted genomic windows per allele, plus positive and negative controls. Two biological replicates were analyzed; the processed table contains the 43 unique mutations represented by the 48 top-significant oligos listed in Extended Data S5 (p<0.01).
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
Basal / Untreated
Synthetic candidate oligos were cloned into the second-generation hSTARR-seq ORI plasmid (Addgene #99296), downstream of the core promoter in the reporter 3′ UTR. Each construct contained a unique 6-bp barcode and a 194-bp genomic sequence context flanked by 15-bp linker regions, for a nominal 230-bp construct. Mutant and wild-type alleles were each represented in five contexts: left_20, left_10, center, right_10, and right_20. PANC-1 cells were transfected with the plasmid pool, poly(A)+ reporter RNA and plasmid DNA were sequenced, and activity was calculated as normalized RNA TPM divided by DNA TPM.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (19 of 19)
| Row | |||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | |||||||||||||||||||
| 2 | |||||||||||||||||||
| 3 | |||||||||||||||||||
| 4 | |||||||||||||||||||
| 5 | |||||||||||||||||||
| 6 | |||||||||||||||||||
| 7 | |||||||||||||||||||
| 8 | |||||||||||||||||||
| 9 | |||||||||||||||||||
| 10 | |||||||||||||||||||
| 11 | |||||||||||||||||||
| 12 | |||||||||||||||||||
| 13 | |||||||||||||||||||
| 14 | |||||||||||||||||||
| 15 | |||||||||||||||||||
| 16 | |||||||||||||||||||
| 17 | |||||||||||||||||||
| 18 | |||||||||||||||||||
| 19 | |||||||||||||||||||
| 20 | |||||||||||||||||||
| 21 | |||||||||||||||||||
| 22 | |||||||||||||||||||
| 23 | |||||||||||||||||||
| 24 | |||||||||||||||||||
| 25 | |||||||||||||||||||
| 26 | |||||||||||||||||||
| 27 | |||||||||||||||||||
| 28 | |||||||||||||||||||
| 29 | |||||||||||||||||||
| 30 | |||||||||||||||||||
| 31 | |||||||||||||||||||
| 32 | |||||||||||||||||||
| 33 | |||||||||||||||||||
| 34 | |||||||||||||||||||
| 35 | |||||||||||||||||||
| 36 | |||||||||||||||||||
| 37 | |||||||||||||||||||
| 38 | |||||||||||||||||||
| 39 | |||||||||||||||||||
| 40 | |||||||||||||||||||
| 41 | |||||||||||||||||||
| 42 | |||||||||||||||||||
| 43 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 19 definitions
- variant_id
- Normalized hg19 variant identifier in the form chromosome:position:reference>alternate; '-' denotes an insertion/deletion allele in the source notation.
- chromosome
- Human chromosome in hg19 notation.
- position_hg19
- 1-based hg19 coordinate of the mutation as reported in the oligo name.
- ref_allele
- Reference allele or deleted sequence from the source oligo name.
- alt_allele
- Alternate allele or inserted sequence from the source oligo name.
- variant_class
- SNV, insertion, deletion, or other indel classification based on the reported alleles.
- annotated_gene_or_element
- Gene or regulatory element label embedded in the Extended Data S5 oligo name.
- top_hit_oligo_ids
- One or more source oligo labels (M1–M48) representing this mutation among the reported top hits; pipe-delimited when multiple windows are listed.
- top_hit_windows
- Source sliding-window labels for the top-hit oligos: L_20, L, M, R, or R_20.
- top_hit_window_offsets_bp
- Nominal mutation-window displacement derived from the source label: -20, -10, 0, 10, or 20 bp.
- top_hit_oligo_count
- Number of top-significant oligos for this mutation listed in Extended Data S5.
- reporter_effect_direction
- Mutant reporter activity direction relative to wild type, mapped from the gain/loss bar colors in Fig. 2e.
- reported_significance
- Significance threshold stated in the Extended Data S5 caption; exact per-oligo p-values were not supplied in the public workbook.
- predicted_tf_motif
- Transcription-factor motif explicitly discussed for this mutation in the paper text/figure; blank when not explicitly reported.
- motif_change
- Explicitly reported motif gain or motif break for the listed transcription factor; blank when not reported.
- reported_variant_mean_log2_fc
- Exact variant-level mean log2 fold change reported in the paper text for the small number of highlighted mutations; blank when not publicly reported.
- reported_variant_p_value
- Exact or inequality p-value for the corresponding text-reported variant-level summary; blank when not publicly reported.
- reported_effect_summary_note
- Scope of any text-reported effect summary, including the number of sliding windows and test description.
- qc_status
- PASS indicates inclusion in the authors’ top-significant-oligo list after the stated sequencing/UMI QC scope.
Quality control
The authors report low-quality read filtering, adapter/linker trimming, 6-bp barcode-based separation of WT, MUT, and control constructs, alignment to the hg19 oligo design, UMI deduplication, and a minimum of 3 unique UMIs per counted construct. The cloned plasmid library had minimum 30× coverage for 98.63% of the 6,082 sequenced constructs with WT and MUT represented. Two biological replicates were retained for the STARR-seq screen. The processed table is further restricted to the authors’ Extended Data S5 list of top-significant oligos (p<0.01), so all retained rows pass the reported screen-level significance/QC scope.
Curation notes
The complete STARR-seq RNA/DNA count matrix, per-oligo activity values, and exact p-values were not deposited publicly; the authors state that STARR-seq data and analysis scripts can be requested from the corresponding authors. Accordingly, table.csv is a transparent derived summary of the public Extended Data S5 top-hit list, not a complete 6,082-oligo assay matrix. The 48 source oligos were grouped into 43 unique mutations, matching the abstract’s 43 prioritized NCMs. Reporter direction was read from Fig. 2e (M1–M37 gain of function; M38–M48 loss of function), and the resulting unique-variant counts are 33 gain and 10 loss. Exact text-reported summaries are retained for chr12:115067012 C>A (mean log2 fold change 3.69 across five sliding windows, p=0.016) and chr3:71123616 G>T (average log2 fold change -1.36 across three significant windows, p<0.05). The paper alternates between 6,080 constructs and 6,082 oligos/constructs in different sections; the table uses the 6,082 value from the library description and QC statement. The source workbook labels the chr12:115067012 C>A region TBX5-AS1, whereas the text describes it as proximal to TBX3; the source label is preserved in the table.