Experiment / E4D55AVR0Standard STARR-seq

ExP STARR-seq 1K×1K enhancer–promoter compatibility screen

Compatibility rules of human enhancer and promoter sequences

A plasmid library paired approximately 1,000 synthetic 264-bp human enhancer sequences with approximately 1,000 264-bp promoter sequences in all combinations, assigning a unique 16-bp plasmid barcode to each construct. The library was transiently transfected into K562 cells in four biological replicates, and reporter RNA was quantified relative to plasmid DNA input.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Episomal self-transcribing reporter using the revised human STARR-seq plasmid. Enhancer and promoter inserts were 264 bp, separated by approximately 340 bp in the construct, and each plasmid carried a random 16-bp barcode adjacent to the enhancer. Four 50-million-cell K562 transfection replicates were sequenced for weighted reporter RNA and plasmid DNA input.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (13 of 13)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 13 definitions
promoter_id
Identifier for the 264-bp promoter sequence in the ExP library; genomic identifiers include hg19 coordinates and controls use scramble labels.
enhancer_id
Identifier for the 264-bp enhancer sequence in the ExP library; genomic identifiers include hg19 coordinates and controls use scramble labels.
barcode_count
Number of plasmid barcodes for this promoter–enhancer pair passing DNA input ≥25 and weighted RNA total ≥1.
dna_input_sum
Sum of DNA input counts across the qualifying barcodes for the pair.
weighted_rna_rep1_sum
Sum of author-supplied weighted reporter RNA counts for biological replicate 1 across qualifying barcodes.
weighted_rna_rep2_sum
Sum of author-supplied weighted reporter RNA counts for biological replicate 2 across qualifying barcodes.
weighted_rna_rep3_sum
Sum of author-supplied weighted reporter RNA counts for biological replicate 3 across qualifying barcodes.
weighted_rna_rep4_sum
Sum of author-supplied weighted reporter RNA counts for biological replicate 4 across qualifying barcodes.
weighted_rna_total_sum
Sum of weighted reporter RNA counts across all four biological replicates and qualifying barcodes.
activity_log2_mean
Mean barcode-level log2 RNA/DNA activity, where RNA is the four-replicate weighted RNA total and DNA is the barcode DNA input count.
activity_log2_sd
Population standard deviation of the qualifying barcode-level log2 RNA/DNA activities for the pair.
activity_log2_min
Minimum qualifying barcode-level log2 RNA/DNA activity for the pair.
activity_log2_max
Maximum qualifying barcode-level log2 RNA/DNA activity for the pair.

Quality control

Author QC: PCR-replicate counts were scaled within biological replicates, barcode–promoter assignments were built from the promoter/barcode dictionary after singleton and ambiguous assignments were removed, and plasmids with fewer than 25 DNA input reads or fewer than 1 weighted RNA read were discarded. Pair-level analysis retained enhancer–promoter pairs with at least two qualifying barcodes. Applying these rules to the deposited GEO file retained 4,647,965 barcodes and 604,270 pairs; the manuscript reports 4,512,907 plasmids used for activity estimation and 604,268 pairs, a two-pair discrepancy attributable to the deposited-file/version definition.

Curation notes

This is the central ExP STARR-seq dataset, not a variant-contrast library: it measures quantitative activity and compatibility of natural genomic and control enhancer/promoter sequences. The processed table is pair-level to avoid treating the barcode technical replicates as independent biological observations; the original barcode-level counts remain in raw_data. Genomic coordinates and controls are reported exactly as named in the GEO file.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.