Experiment / E5H5G6MHIStandard STARR-seq

Initial 194-bp candidate enhancer tile STARR-seq screen in HepG2

Functional characterization of enhancer evolution in the primate lineage

An episomal STARR-seq screen tested 10,544 synthesized 194-bp tiles from 1,015 candidate hominoid-specific liver enhancers together with dinucleotide-shuffled negative controls. Three biological transfections in HepG2 cells were quantified as normalized reporter-RNA/input-DNA log2 enrichment scores.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated; 1 ng/mL puromycin selection for 24 h after transfection

Episomal human STARR-seq vector with candidate sequences cloned into the reporter 3′ UTR. A 244 K Agilent array supplied 194-bp tiles and 800 dinucleotide-shuffled controls; 5 µg library plus 2.5 µg puromycin plasmid was transfected with Lipofectamine 3000 into three approximately 1.5-million-cell HepG2 dishes. DNA and RNA were collected 48 h post-transfection, sequenced as paired DNA/RNA libraries on an Illumina NextSeq 500/550 with PE150 reads, and scored as normalized RNA/DNA enrichment.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (13 of 13)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 13 definitions
element_id
Stable package identifier assigned in decompressed source-row order.
source_element_id
Original element identifier from the GEO score table; coordinates or the literal Negative label.
element_type
Parsed source class: Candidate tile or Negative control.
coordinate_hg19
Original hg19-style coordinate string; blank for Negative controls.
chromosome
Chromosome parsed from coordinate_hg19.
start
Start boundary parsed from the deposited coordinate.
end
End boundary parsed from the deposited coordinate.
length_bp
end minus start in the source boundary convention; candidate tiles are 194 bp and Negative controls are blank.
activity_score_log2_rna_dna
Author-deposited log2 normalized reporter-RNA/input-DNA enrichment score.
score_above_author_cutoff
TRUE when the deposited score is greater than 1, the paper's active-tile threshold.
qc_pass
TRUE for rows retained after package-level identifier, score, and coordinate integrity QC.
source_row
1-based row number in the decompressed GEO file, including its header as row 1.
source_file
Relative path to the compressed source table in raw_data.

Quality control

Author QC was retained: reads were aligned to the input library with BWA-MEM; normalized RNA/DNA scores used a hard DNA-read cutoff >10 and excluded ratios with zero RNA reads; the paper defined active tiles as log2 enrichment >1. Package integrity QC additionally required a nonempty element identifier, a finite numeric score, and either a valid positive hg19-style coordinate or the deposited Negative-control label. All 6,859 deposited rows passed these checks and are retained; the score-above-cutoff flag is TRUE for 697 rows, matching the paper's reported active-tile count, while three of those rows remain visibly labeled Negative controls.

Curation notes

Source: GSE113978_Tiling_scores.tsv.gz and Supplementary Table S1. The source table contains 6,859 finite scores from the larger 10,544-element library after the author-described DNA-coverage and zero-RNA exclusions, including 124 rows labeled Negative. Eleven coordinate groups are duplicated in the source; package element_id values preserve those independent rows instead of collapsing them. Coordinates and length_bp retain the source's BED-like end-minus-start convention. No raw sequence reads were packaged.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.