Experiment / E2IJWITKCATAC-STARR-seq

HiDRA genome-wide ATAC-STARR-seq in GM12878

High-resolution genome-wide functional dissection of transcriptional regulatory regions and nucleotides in human

HiDRA captured accessible human genomic fragments from GM12878 cells, cloned them into an episomal STARR-seq reporter, and transfected the same lymphoblastoid cell line in five biological replicates. Matched plasmid-input and RNA-output sequencing quantified regulatory activity across overlapping fragment groups, with SHARPR-RE used by the authors to infer high-resolution driver elements.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

The library was made from Tn5-accessible chromatin in GM12878, size-selected to approximately 150–500 nt, and cloned into the 3′ UTR of the non-integrating pSTARR-seq_human reporter (Addgene plasmid #71509). Five independent transfections of approximately 120–130 million GM12878 cells were harvested 24 h later for RNA; five matched plasmid-input samples were sequenced alongside the RNA samples. The public GEO matrix contains group-summed counts for P1–P5 and R1–R5 after 75% reciprocal-overlap grouping. The same library was also used for auxiliary allele-specific analysis of heterozygous GM12878 variants, and those annotations are joined where the public GEO fragment-group identifiers match.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (35 of 35)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 35 definitions
fragment_group_id
GEO fragment-group identifier in the form chromosome:start-end_num_unique_fragments.
chromosome
Chromosome of the fragment group on hg19.
start_hg19
0-based inclusive hg19 start coordinate from the GEO/BED-style source.
end_hg19
0-based half-open hg19 end coordinate from the GEO/BED-style source.
length_bp
Fragment-group span in base pairs, calculated as end_hg19 minus start_hg19.
num_unique_fragments
Number of unique constituent fragments encoded in the GEO identifier suffix.
dna_count_p1
Raw plasmid-input read count for biological replicate P1.
dna_count_p2
Raw plasmid-input read count for biological replicate P2.
dna_count_p3
Raw plasmid-input read count for biological replicate P3.
dna_count_p4
Raw plasmid-input read count for biological replicate P4.
dna_count_p5
Raw plasmid-input read count for biological replicate P5.
rna_count_r1
Raw reporter-RNA read count for biological replicate R1.
rna_count_r2
Raw reporter-RNA read count for biological replicate R2.
rna_count_r3
Raw reporter-RNA read count for biological replicate R3.
rna_count_r4
Raw reporter-RNA read count for biological replicate R4.
rna_count_r5
Raw reporter-RNA read count for biological replicate R5.
dna_total_count
Sum of plasmid-input counts across P1–P5.
rna_total_count
Sum of reporter-RNA counts across R1–R5.
dna_mean_cpm
Mean plasmid-input counts per million across P1–P5 using source-library totals.
rna_mean_cpm
Mean reporter-RNA counts per million across R1–R5 using source-library totals.
log2_activity_rep1
Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 1.
log2_activity_rep2
Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 2.
log2_activity_rep3
Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 3.
log2_activity_rep4
Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 4.
log2_activity_rep5
Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 5.
mean_log2_activity
Arithmetic mean of the five replicate-level log2 activity scores.
median_log2_activity
Median of the five replicate-level log2 activity scores.
activity_sd_log2
Sample standard deviation of the five replicate-level log2 activity scores.
active_region_coord
Published Supplementary Data 1 active HiDRA region(s) overlapped by the fragment group, in hg19 coordinates.
tiled_region_coord
Published Supplementary Data 2 SHARPR-RE tiled region(s) overlapped by the fragment group, in hg19 coordinates.
driver_element_coords
Published Supplementary Data 3 high-resolution driver element(s) overlapped by the fragment group, in hg19 coordinates.
allele_specific_snp_count
Number of distinct SNP positions in the GEO allele-specific count file associated with this fragment group.
allele_specific_snp_annotations
Allele-specific SNP annotations associated with the group, formatted as chromosome:position:reference>alternate and separated by semicolons.
significant_allele_specific_snp_raw_pvalues
Supplementary Data 4 allele-specific SNP positions with the reported unadjusted p-value <0.05, formatted as chromosome:position=p-value and separated by semicolons.
min_significant_allele_specific_snp_raw_pvalue
Minimum reported unadjusted p-value among significant allele-specific SNP annotations for the group; blank when none are present.

Quality control

The authors aligned reads to hg19, retained aligned fragments, removed chrM reads, required mapping quality ≥30, removed ENCODE hg19 blacklist intervals, and retained fragments 100–600 nt long before grouping fragments by 75% mutual overlap and summing counts. Paper-level activity calls used DESeq2 with fragment groups split into 100-nt bins and an FDR <0.05 threshold. Package QC retained 296,223 of 313,170 public fragment-group rows that overlapped one of the 66,254 published active HiDRA regions and had aggregate plasmid DNA count >0 and aggregate RNA count >0 across the five replicates; 16,947 rows lacking one of those aggregate signals were excluded. The table does not recreate the paper’s unavailable per-group DESeq2 p-values/FDR; it reports transparent package-derived CPM-normalized activity metrics using a 0.1 CPM pseudocount.

Curation notes

This study contributes one MPRA-family experiment: a non-integrating, self-transcribing ATAC-STARR-seq/HiDRA assay in the GM12878 lymphoblastoid cell line (Cellosaurus CVCL:7526). The authors reported five RNA transfection replicates and five plasmid-input replicates, plus an auxiliary allele-specific analysis of the same library rather than a separate biological condition. The public GEO count release has 7,563,783 fragment-group rows, while the article narrative summarizes approximately 7.1 million groups; the released counts were not altered. The processed table is intentionally limited to QC-passing groups overlapping the paper’s active-region BED file and retains published tiled and driver-element overlaps. Because the public count matrix does not include the paper’s per-group DESeq2 statistics or exact active-group calls, the activity columns are package-derived from library-size CPM and are not presented as DESeq2 effect sizes. Raw GEO counts, allele-specific files, and supplementary BED/text files are preserved under raw_data; raw sequencing reads and GEO bigWig tracks were omitted to keep the study package compact.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.