Experiment / E1C44NBMBWhole-Genome STARR-seq (WHG-STARR-seq)

Pooled whole-genome STARR-seq in LNCaP cells

Functional assessment of human enhancer activities using whole-genome STARR-sequencing

An episomal WHG-STARR-seq library of randomly sheared human genomic fragments was tested in the human prostate cancer cell line LNCaP using the SCP1 minimal promoter. Two biological replicates were pooled for MACS2 peak calling against the input plasmid library; the packaged table contains 94,527 unique active enhancer peaks passing the reported q-value threshold.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Episomal self-transcribing reporter assay using 350–650 bp size-selected fragments from random human genomic DNA, with an average fragment length of approximately 500 bp. Inserts were cloned downstream of the SCP1 minimal promoter and into the reporter 3' UTR; the input plasmid library served as the DNA baseline and the poly(A)-selected self-transcribed RNA output was the activity readout. Paired-end 2 × 100 bp sequencing was aligned to hg19, and the authors used uniquely mapped distinct fragments, pooled across 16 indexing libraries per replicate, for combined two-replicate MACS2 peak calling.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (16 of 16)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 16 definitions
element_id
Unique WHG-STARR-seq peak identifier from the GEO processed file.
chromosome
Human hg19 chromosome name.
start_hg19_0based
BED interval start coordinate on hg19, zero-based and inclusive.
end_hg19_0based_exclusive
BED interval end coordinate on hg19, zero-based and exclusive.
length_bp
Peak interval length in base pairs, calculated as end minus start.
macs2_score
Integer score from the deposited MACS2 narrowPeak-style record.
enhancer_activity_fold_enrichment
MACS2 signal/enrichment value; the paper uses this enrichment score as WHG-STARR-seq enhancer activity.
p_value_neg_log10
MACS2 peak significance reported as -log10(p-value).
q_value_neg_log10
MACS2 multiple-testing significance reported as -log10(q-value).
p_value
P-value derived from the deposited -log10(p-value) field as 10 raised to the negative reported value.
q_value
Q-value derived from the deposited -log10(q-value) field as 10 raised to the negative reported value.
summit_offset_bp
MACS2 summit offset from the peak interval start in base pairs.
chromatin_context
Active enhancer context classified by overlap with the LNCaP DNase I sites: open for an overlap and closed for no overlap.
dnase_peak_count
Number of unique deposited DNase I peaks overlapping the WHG-STARR-seq peak; zero for closed peaks.
dnase_peak_ids
Semicolon-separated identifiers of overlapping deposited DNase I peaks; empty for closed peaks.
source_file
GEO processed source file used for the row, indicating the open/closed enhancer partition.

Quality control

The authors filtered multi-mapping reads, collapsed fragments with identical inferred start and end positions using Picard, pooled distinct fragments from the 16 indexing libraries for each replicate, combined the two biological replicates, and called active peaks with MACS2 at q-value < 0.05 using the input plasmid library as background. For this package, records were additionally required to have valid genomic intervals, complete numeric peak fields, and q-value <= 0.05; repeated STARR–DNase overlap rows were collapsed to one row per unique STARR peak. No records were removed by these package-level checks.

Curation notes

This is a region-focused, non-allelic assay: the library was made from random sheared genomic DNA and does not provide rsIDs or reference/alternative allele contrasts. GEO's enhancer_overlap_dnase file contains one row per overlap, so its 12,971 rows were aggregated to 11,980 unique open STARR peaks; the 82,547 non-overlap rows were already unique. The final table therefore has 94,527 rows, matching the paper's reported active enhancer count, with 11,980 open and 82,547 closed peaks. The two biological STARR-seq replicates are represented as one pooled experiment because the authors pooled them for final peak calling. The accompanying ATAC-seq and RNA-seq files are retained in raw_data as study context and provenance, not merged into the MPRA activity table.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.