HiDRA genome-wide ATAC-STARR-seq in GM12878
High-resolution genome-wide functional dissection of transcriptional regulatory regions and nucleotides in humanHiDRA captured accessible human genomic fragments from GM12878 cells, cloned them into an episomal STARR-seq reporter, and transfected the same lymphoblastoid cell line in five biological replicates. Matched plasmid-input and RNA-output sequencing quantified regulatory activity across overlapping fragment groups, with SHARPR-RE used by the authors to infer high-resolution driver elements.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
Basal / Untreated
The library was made from Tn5-accessible chromatin in GM12878, size-selected to approximately 150–500 nt, and cloned into the 3′ UTR of the non-integrating pSTARR-seq_human reporter (Addgene plasmid #71509). Five independent transfections of approximately 120–130 million GM12878 cells were harvested 24 h later for RNA; five matched plasmid-input samples were sequenced alongside the RNA samples. The public GEO matrix contains group-summed counts for P1–P5 and R1–R5 after 75% reciprocal-overlap grouping. The same library was also used for auxiliary allele-specific analysis of heterozygous GM12878 variants, and those annotations are joined where the public GEO fragment-group identifiers match.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (35 of 35)
| Row | |||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | |||||||||||||||||||||||||||||||||||
| 2 | |||||||||||||||||||||||||||||||||||
| 3 | |||||||||||||||||||||||||||||||||||
| 4 | |||||||||||||||||||||||||||||||||||
| 5 | |||||||||||||||||||||||||||||||||||
| 6 | |||||||||||||||||||||||||||||||||||
| 7 | |||||||||||||||||||||||||||||||||||
| 8 | |||||||||||||||||||||||||||||||||||
| 9 | |||||||||||||||||||||||||||||||||||
| 10 | |||||||||||||||||||||||||||||||||||
| 11 | |||||||||||||||||||||||||||||||||||
| 12 | |||||||||||||||||||||||||||||||||||
| 13 | |||||||||||||||||||||||||||||||||||
| 14 | |||||||||||||||||||||||||||||||||||
| 15 | |||||||||||||||||||||||||||||||||||
| 16 | |||||||||||||||||||||||||||||||||||
| 17 | |||||||||||||||||||||||||||||||||||
| 18 | |||||||||||||||||||||||||||||||||||
| 19 | |||||||||||||||||||||||||||||||||||
| 20 | |||||||||||||||||||||||||||||||||||
| 21 | |||||||||||||||||||||||||||||||||||
| 22 | |||||||||||||||||||||||||||||||||||
| 23 | |||||||||||||||||||||||||||||||||||
| 24 | |||||||||||||||||||||||||||||||||||
| 25 | |||||||||||||||||||||||||||||||||||
| 26 | |||||||||||||||||||||||||||||||||||
| 27 | |||||||||||||||||||||||||||||||||||
| 28 | |||||||||||||||||||||||||||||||||||
| 29 | |||||||||||||||||||||||||||||||||||
| 30 | |||||||||||||||||||||||||||||||||||
| 31 | |||||||||||||||||||||||||||||||||||
| 32 | |||||||||||||||||||||||||||||||||||
| 33 | |||||||||||||||||||||||||||||||||||
| 34 | |||||||||||||||||||||||||||||||||||
| 35 | |||||||||||||||||||||||||||||||||||
| 36 | |||||||||||||||||||||||||||||||||||
| 37 | |||||||||||||||||||||||||||||||||||
| 38 | |||||||||||||||||||||||||||||||||||
| 39 | |||||||||||||||||||||||||||||||||||
| 40 | |||||||||||||||||||||||||||||||||||
| 41 | |||||||||||||||||||||||||||||||||||
| 42 | |||||||||||||||||||||||||||||||||||
| 43 | |||||||||||||||||||||||||||||||||||
| 44 | |||||||||||||||||||||||||||||||||||
| 45 | |||||||||||||||||||||||||||||||||||
| 46 | |||||||||||||||||||||||||||||||||||
| 47 | |||||||||||||||||||||||||||||||||||
| 48 | |||||||||||||||||||||||||||||||||||
| 49 | |||||||||||||||||||||||||||||||||||
| 50 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 35 definitions
- fragment_group_id
- GEO fragment-group identifier in the form chromosome:start-end_num_unique_fragments.
- chromosome
- Chromosome of the fragment group on hg19.
- start_hg19
- 0-based inclusive hg19 start coordinate from the GEO/BED-style source.
- end_hg19
- 0-based half-open hg19 end coordinate from the GEO/BED-style source.
- length_bp
- Fragment-group span in base pairs, calculated as end_hg19 minus start_hg19.
- num_unique_fragments
- Number of unique constituent fragments encoded in the GEO identifier suffix.
- dna_count_p1
- Raw plasmid-input read count for biological replicate P1.
- dna_count_p2
- Raw plasmid-input read count for biological replicate P2.
- dna_count_p3
- Raw plasmid-input read count for biological replicate P3.
- dna_count_p4
- Raw plasmid-input read count for biological replicate P4.
- dna_count_p5
- Raw plasmid-input read count for biological replicate P5.
- rna_count_r1
- Raw reporter-RNA read count for biological replicate R1.
- rna_count_r2
- Raw reporter-RNA read count for biological replicate R2.
- rna_count_r3
- Raw reporter-RNA read count for biological replicate R3.
- rna_count_r4
- Raw reporter-RNA read count for biological replicate R4.
- rna_count_r5
- Raw reporter-RNA read count for biological replicate R5.
- dna_total_count
- Sum of plasmid-input counts across P1–P5.
- rna_total_count
- Sum of reporter-RNA counts across R1–R5.
- dna_mean_cpm
- Mean plasmid-input counts per million across P1–P5 using source-library totals.
- rna_mean_cpm
- Mean reporter-RNA counts per million across R1–R5 using source-library totals.
- log2_activity_rep1
- Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 1.
- log2_activity_rep2
- Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 2.
- log2_activity_rep3
- Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 3.
- log2_activity_rep4
- Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 4.
- log2_activity_rep5
- Package-derived log2((RNA CPM + 0.1)/(DNA CPM + 0.1)) for paired replicate 5.
- mean_log2_activity
- Arithmetic mean of the five replicate-level log2 activity scores.
- median_log2_activity
- Median of the five replicate-level log2 activity scores.
- activity_sd_log2
- Sample standard deviation of the five replicate-level log2 activity scores.
- active_region_coord
- Published Supplementary Data 1 active HiDRA region(s) overlapped by the fragment group, in hg19 coordinates.
- tiled_region_coord
- Published Supplementary Data 2 SHARPR-RE tiled region(s) overlapped by the fragment group, in hg19 coordinates.
- driver_element_coords
- Published Supplementary Data 3 high-resolution driver element(s) overlapped by the fragment group, in hg19 coordinates.
- allele_specific_snp_count
- Number of distinct SNP positions in the GEO allele-specific count file associated with this fragment group.
- allele_specific_snp_annotations
- Allele-specific SNP annotations associated with the group, formatted as chromosome:position:reference>alternate and separated by semicolons.
- significant_allele_specific_snp_raw_pvalues
- Supplementary Data 4 allele-specific SNP positions with the reported unadjusted p-value <0.05, formatted as chromosome:position=p-value and separated by semicolons.
- min_significant_allele_specific_snp_raw_pvalue
- Minimum reported unadjusted p-value among significant allele-specific SNP annotations for the group; blank when none are present.
Quality control
The authors aligned reads to hg19, retained aligned fragments, removed chrM reads, required mapping quality ≥30, removed ENCODE hg19 blacklist intervals, and retained fragments 100–600 nt long before grouping fragments by 75% mutual overlap and summing counts. Paper-level activity calls used DESeq2 with fragment groups split into 100-nt bins and an FDR <0.05 threshold. Package QC retained 296,223 of 313,170 public fragment-group rows that overlapped one of the 66,254 published active HiDRA regions and had aggregate plasmid DNA count >0 and aggregate RNA count >0 across the five replicates; 16,947 rows lacking one of those aggregate signals were excluded. The table does not recreate the paper’s unavailable per-group DESeq2 p-values/FDR; it reports transparent package-derived CPM-normalized activity metrics using a 0.1 CPM pseudocount.
Curation notes
This study contributes one MPRA-family experiment: a non-integrating, self-transcribing ATAC-STARR-seq/HiDRA assay in the GM12878 lymphoblastoid cell line (Cellosaurus CVCL:7526). The authors reported five RNA transfection replicates and five plasmid-input replicates, plus an auxiliary allele-specific analysis of the same library rather than a separate biological condition. The public GEO count release has 7,563,783 fragment-group rows, while the article narrative summarizes approximately 7.1 million groups; the released counts were not altered. The processed table is intentionally limited to QC-passing groups overlapping the paper’s active-region BED file and retains published tiled and driver-element overlaps. Because the public count matrix does not include the paper’s per-group DESeq2 statistics or exact active-group calls, the activity columns are package-derived from library-size CPM and are not presented as DESeq2 effect sizes. Raw GEO counts, allele-specific files, and supplementary BED/text files are preserved under raw_data; raw sequencing reads and GEO bigWig tracks were omitted to keep the study package compact.