Experiment / E5NPBAJE6Standard STARR-seq

Tiling STARR-seq enhancer activity in HCT116 cells

Systematic discovery and functional dissection of enhancers needed for cancer cell fitness and proliferation

An episomal hSTARR-seq library comprising 12,833 190-bp fragments tiled across HCT116 essential-enhancer windows, 60 positive-control oligos for known MYC/SV40 enhancers, and 2,095 yeast ORF negative controls was electroporated into HCT116 cells in triplicate. Matched plasmid DNA and poly(A)+ reporter-RNA counts from Table S3 were used to quantify per-oligo enhancer activity.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

The authors cloned 190-bp oligos with constant flanking sequences into the episomal hSTARR-seq_ORI reporter (Addgene plasmid #99296), electroporated HCT116 cells with TBK1/IKK and PKR inhibitors, harvested cells 48 h later, and sequenced plasmid DNA and poly(A)+ reporter RNA. Unique molecular identifiers were incorporated during library preparation; sequencing used 2 x 100-bp paired-end reads on an Illumina NextSeq 2000, with three biological transfection replicates.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (25 of 25)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 25 definitions
oligo_id
Oligonucleotide identifier from Table S3; genomic fragments use hg38 coordinate labels and controls use yeast ORF or SV40 labels.
element_class
Class inferred from the identifier: human_genomic_fragment, genomic_negative_control, yeast_orf_negative_control, or SV40_positive_control.
sequence_role
Role inferred from the identifier: negative_control, positive_control, or genomic_test_or_known_positive.
chromosome
Chromosome parsed from a genomic coordinate identifier; blank for synthetic yeast and SV40 controls.
start_hg38
Start coordinate parsed from the genomic oligo identifier, using the coordinate convention reported in Table S3.
end_hg38
End coordinate parsed from the genomic oligo identifier, using the coordinate convention reported in Table S3.
sequence
The 190-bp genomic or control sequence from the Table S3 oligo-sequence worksheet.
sequence_length
Length of the oligo sequence in base pairs.
plasmid_dna_count
Raw plasmid-DNA input read count from the HCT116 raw-count worksheet.
rna_rep1_count
Raw reporter-RNA read count for HCT116 biological replicate 1.
rna_rep2_count
Raw reporter-RNA read count for HCT116 biological replicate 2.
rna_rep3_count
Raw reporter-RNA read count for HCT116 biological replicate 3.
mean_rna_count
Arithmetic mean of the three raw HCT116 reporter-RNA counts.
dna_cpm
Plasmid-DNA count normalized to counts per million over the retained oligo set.
rna_rep1_cpm
Replicate-1 reporter-RNA count normalized to counts per million over the retained oligo set.
rna_rep2_cpm
Replicate-2 reporter-RNA count normalized to counts per million over the retained oligo set.
rna_rep3_cpm
Replicate-3 reporter-RNA count normalized to counts per million over the retained oligo set.
mean_rna_cpm
Arithmetic mean of the three normalized reporter-RNA CPM values.
log2_activity_rep1
Library-size-normalized log2(RNA CPM / plasmid-DNA CPM) for replicate 1.
log2_activity_rep2
Library-size-normalized log2(RNA CPM / plasmid-DNA CPM) for replicate 2.
log2_activity_rep3
Library-size-normalized log2(RNA CPM / plasmid-DNA CPM) for replicate 3.
mean_log2_activity
Mean of the three replicate log2 RNA/DNA activity values; larger values indicate greater reporter activity.
activity_sd
Population standard deviation of the three replicate log2 activity values.
activity_rank
Descending rank among retained oligos by mean_log2_activity, with 1 as the highest activity.
source_table
The Table S3 worksheets used to assemble the row: oligo sequence and raw counts HCT116.

Quality control

The paper aligned STARR-seq reads to the oligo library, removed oligos with plasmid-DNA counts <25 or RNA counts <5 in at least three biological replicate libraries, estimated dispersion from yeast ORF controls, and used edgeR negative-binomial regression with Benjamini-Hochberg adjustment and a 5% empirical FDR threshold calibrated to enriched yeast controls. The paper reports Pearson correlation >0.9 for RNA/DNA ratios across the three biological replicates. Package QC retained the 14,852 unique oligos in the Table S3 raw-count worksheet: all had finite nonnegative counts, plasmid DNA >=25 (minimum 32), and all three RNA replicate counts >=5 (minimum 6). The 14,988-row sequence worksheet was joined by oligo ID; sequence rows without a corresponding count record were excluded. The generated activity columns are descriptive library-size-normalized CPM and log2(RNA CPM/DNA CPM) values; author edgeR p-values and calls were not recreated because they are not supplied in Table S3.

Curation notes

The Cellosaurus accession for the parental HCT116 line is CVCL_0291 (represented here in the requested CURIE form CVCL:0291). The source workbook supplies raw counts only for HCT116; the paper reports a matched K562 STARR-seq comparison in Figure 4H but does not provide a K562 count worksheet or separate downloadable result table. K562 values were therefore not inferred. The human_genomic_fragment class includes candidate enhancer tiles and unlabeled known-MYC positive-control tiles because Table S3 does not distinguish those subsets in its identifiers.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.