Experiment / E5J4VPAYMStandard STARR-seq

HEK293T STARR-seq allelic enhancer activity screen

Systematic analysis of binding of transcription factors to noncoding variants

A pooled human STARR-seq library tested 190-bp genomic fragments containing alleles of 3,943 noncoding SNPs in HEK293T cells, together with 37 known enhancer controls and 2,998 yeast ORF negative controls. Three biological replicates were generated, with two technical sequencing libraries per replicate.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Episomal human STARR-seq plasmids were transfected into HEK293T cells with Fugene HD and harvested 48 hours later. The library used SNP-centered and ATAC-summit-centered 190-bp genomic inserts with constant amplification/cloning flanks, plus known enhancer and yeast ORF controls; poly(A)+ reporter RNA was reverse-transcribed, amplified, and sequenced. GEO labels the two technical library types as ssIII and AB; these were summed within each of three biological replicates before processing.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (42 of 42)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 42 definitions
oligo_id
Exact oligo identifier from the GEO count matrix and Supplementary Table 5.
element_id
Context-specific element identifier; variant contexts distinguish SNP-centered and summit-centered inserts.
variant_id
Chromosome and 1-based SNP position in chr:position form; blank for controls.
rs_id
dbSNP rs identifier recovered by matching the coordinate to Supplementary Table 1.
element_type
SNP-containing genomic fragment, known enhancer positive control, or yeast ORF negative control.
design_set
Supplementary Table 5 library set: snpcenter, summitcenter, positive, or yeast.
chromosome
Chromosome or yeast ORF identifier from Supplementary Table 5.
position
Source position; human variant positions are 1-based, while yeast control positions retain the source midpoint.
fragment_start
Insert start coordinate parsed from the oligo identifier where supplied; blank for SNP-centered inserts without explicit bounds.
fragment_end
Insert end coordinate parsed from the oligo identifier where supplied; blank for SNP-centered inserts without explicit bounds.
allele
Allele represented by the oligo; NONE for non-allelic controls.
reference_allele
Reference allele from the author pbSNP or paSNP annotation when available.
alternative_alleles
Semicolon-separated alternative allele(s) from the author pbSNP or paSNP annotation.
allele_role
Reference, alternative, other, or unoriented relative to the available author allele annotation.
reported_pbSNP
TRUE when the SNP is listed in Supplementary Table 3 as a preferential-binding SNP; blank for controls.
pbSNP_affected_TFs
Semicolon-separated transcription factors listed as affected for the pbSNP in Supplementary Table 3.
reported_paSNP
TRUE when the coordinate is listed in Supplementary Table 6 as preferentially active in HEK293; blank for controls.
paSNP_reference_allele
Reference allele in the cell-specific Supplementary Table 6 paSNP call.
paSNP_alternative_alleles
Semicolon-separated alternative allele(s) in the cell-specific Supplementary Table 6 paSNP call.
input_count
Raw count in the synthesized plasmid input library.
input_cpm
Input count normalized to total input-library counts as counts per million.
rna_count_rep1
RNA reporter count for biological replicate 1 after summing ssIII and AB technical libraries.
rna_count_rep2
RNA reporter count for biological replicate 2 after summing ssIII and AB technical libraries.
rna_count_rep3
RNA reporter count for biological replicate 3 after summing ssIII and AB technical libraries.
rna_cpm_rep1
Replicate 1 RNA count normalized to the merged HEK293T RNA library total as counts per million.
rna_cpm_rep2
Replicate 2 RNA count normalized to the merged HEK293T RNA library total as counts per million.
rna_cpm_rep3
Replicate 3 RNA count normalized to the merged HEK293T RNA library total as counts per million.
log2_activity_rep1
Library-size-normalized log2 RNA/input activity for replicate 1, using a 0.5-count pseudocount.
log2_activity_rep2
Library-size-normalized log2 RNA/input activity for replicate 2, using a 0.5-count pseudocount.
log2_activity_rep3
Library-size-normalized log2 RNA/input activity for replicate 3, using a 0.5-count pseudocount.
mean_log2_activity
Mean of the three replicate log2 RNA/input activity scores.
sd_log2_activity
Sample standard deviation of the three replicate log2 activity scores.
mean_activity_fold
2 raised to mean_log2_activity; normalized RNA/input activity on a fold scale.
n_alleles_designed
Number of designed oligo alleles in the same context-specific element before QC.
n_alleles_retained
Number of allele oligos retained for the same context-specific element after QC.
allele_activity_range_log2
Difference between the highest and lowest retained allele mean log2 activity in the element.
most_active_allele
Allele with the highest retained mean log2 activity in the element.
least_active_allele
Allele with the lowest retained mean log2 activity in the element.
allele_effect_vs_reference_log2
Agent-derived mean alternative-minus-reference log2 activity across three replicates for this alternative allele.
agent_paired_t_pvalue_vs_reference
Agent-derived two-sided paired t-test p-value across three replicate activity differences versus the reference allele; not the paper's limma p-value.
agent_paired_t_fdr_vs_reference
Benjamini-Hochberg FDR for the agent-derived paired tests across retained HEK293T element/allele comparisons.
qc_pass
TRUE for every row retained after the stated input and merged-replicate coverage filter; failed rows are omitted.

Quality control

The paper's STARR-seq QC required input coverage >25 reads and >5 reads in at least three libraries after technical replicate handling, and used yeast oligos to estimate common dispersion for edgeR enrichment testing. Here ssIII and AB counts were summed within each biological replicate, then the stated filter was applied: 14,925 of 14,996 oligos passed and 71 were removed (57 failed the input threshold; 53 failed the coverage threshold; 71 failed at least one threshold). The paper's reported paSNP calls (Supplementary Table 6; author FDR <0.01 from paired limma testing) are retained as annotations. The processed table also contains an agent-derived exploratory paired t-test/FDR across three replicate log2 activity scores; these values are not the authors' edgeR/limma statistics and were not used for filtering.

Curation notes

This child experiment packages the paper's genuine STARR-seq validation rather than the study's primary SNP-SELEX binding measurements. HEK293T was resolved to Cellosaurus CVCL:0063; the source GEO sample titles use HEK293T while Supplementary Table 6 uses HEK293. Supplementary Table 6 contains 206 rows representing 205 distinct coordinates after collapsing duplicate allele entries; all 205 coordinates remain represented in the QC-passing table. The paSNP flag is repeated for every retained allele/context row at a reported coordinate, including both SNP-centered and summit-centered constructs when present. The processed table omits the long insert sequence to keep the analysis table compact; the exact sequences remain in raw_data/Supplementary_Table_5_STARR_seq_oligos.xlsx and its CSV conversion. The calculated activity scores are library-size-normalized summaries of the supplied counts and should not be treated as a reimplementation of the authors' edgeR model output.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.