Experiment / E2KXEB5RVTargeted / Cap-STARR-seq

NSC ChIP-STARR-seq: NSC-derived HIST and TF libraries in neural stem cells

BRAIN-MAGNET: A novel functional genomics atlas coupled with convolutional neural networks facilitates clinical interpretation of disease relevant variants in non-coding regulatory elements

Pooled ChIP-STARR-seq plasmid libraries were generated from H9-derived neural stem cell chromatin: the HIST pool combined H3K27ac and H3K4me1 ChIP DNA, and the TF pool combined SOX2 and YY1 ChIP DNA. The pooled libraries were transfected into H9-derived neural stem cells, and enhancer activity was measured for 148,198 genomic scaffold regions as normalized STARR RNA/plasmid DNA activity.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Adapted ChIP-STARR-seq episomal self-transcribing reporter assay; target enrichment was performed by ChIP of NSC chromatin rather than capture hybridization. Genomic fragments enriched by NSC ChIP were cloned downstream of a minimal-promoter-driven GFP and upstream of the polyadenylation signal in a STARR-seq plasmid. Five independent transfections per plasmid pool were pooled 24 hours later; RNA and plasmid DNA sequencing were used to calculate activity. The paper reports four indexed RNA replicates per HIST or TF pool and two plasmid-library DNA replicates per pool.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (16 of 16)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 16 definitions
element_id
Published NCREs_location identifier in chr:start-end form.
chrom
Chromosome or reference contig parsed from element_id.
start
Reported scaffold start coordinate parsed from element_id.
end
Reported scaffold end coordinate parsed from element_id.
length_bp
Scaffold length in base pairs, calculated as end minus start.
activity_log2_rpp1
Published NSC ChIP-STARR-seq activity, log2(avg.NSC.RPP+1), where RPP is the normalized RNA/plasmid read ratio.
activity_category_10
Published NSC activity decile from 10_Categories; Category_10 is the highest-activity decile.
activity_category_5
Published NSC activity quintile from 5_Categories; Category_5 is the highest-activity quintile.
comparison_esc_activity_log2_rpp1
Published comparative ChIP-STARR-seq activity after transfection of the same NSC-derived library into ESCs, log2(avg.ESC.RPP+1).
comparison_esc_activity_category_10
Published comparative ESC activity decile for the same scaffold; Category_10 is the highest-activity decile.
comparison_esc_activity_category_5
Published comparative ESC activity quintile for the same scaffold; Category_5 is the highest-activity quintile.
target_gene_ids
Unique condition-specific target Ensembl gene IDs from Table S2, pipe-delimited.
target_gene_names
Unique condition-specific target gene names from Table S2, pipe-delimited.
target_gene_types
Unique target gene biotypes from Table S2, pipe-delimited.
omim_phenotypes
Unique OMIM phenotype annotations associated with the target genes, pipe-delimited.
hpo_terms
Unique HPO term CURIEs associated with the target genes, pipe-delimited.

Quality control

Study QC: NSC ChIP was performed in duplicate and ChIP-seq replicates showed Pearson correlation >0.88; independent indexed STARR RNA technical replicates showed Pearson correlation >0.96 and were merged for downstream analysis. Reads were adapter-trimmed, aligned to GRCh38/hg38 with Bowtie2 using --very-sensitive, and only properly and uniquely mapped reads with MAPQ >=30 were retained; ChIP duplicates were removed with Picard. Study scaffold activity was computed after removing low-coverage regions with fewer than 20 reads in at least two samples, using normalized RNA/plasmid read ratios averaged across replicates. Package QC retained one row per unique published NCRE location, requiring finite NSC and ESC log2(avg.RPP+1) values, parseable chromosome coordinates, and a 500-1000 bp scaffold length. All 148,198 unique scaffolds passed these package filters.

Curation notes

The source Table S2 contains repeated rows for a scaffold when it has multiple target-gene/HPO annotations. The processed table collapses exact duplicate coordinates and combines unique annotations with pipes; activity values and NSC categories are taken from the NSC sheet, while the ESC activity/category fields are retained for comparison. The paper removed a small number of patch sequences before CNN training, but Table S2 does not expose a patch flag; this package represents the published measured activity atlas and does not silently remove those rows. Individual low-throughput validation constructs and BRAIN-MAGNET nucleotide contribution scores were not substituted for the atlas activity table. Biosample CVCL:IU37 is the Cellosaurus H9-NSC entry.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.