Experiment / E2438D8A1AAV-MPRA / in vivo MPRA

STAR408 scAAV MPRA of human enhancer candidates in early postnatal mouse brain

Parallel functional testing identifies enhancers active in early postnatal mouse brain

The main STAR408 screen tested 408 approximately 900 bp human candidate regulatory sequences from GWAS, LD, FBDHS, and PutEnh design groups in P7 mouse forebrain after P0 scAAV delivery. Four biological replicate DNA and RNA libraries were used to quantify enhancer activity by normalized RNA/DNA ratios and a GC-adjusted residual model; the packaged table retains the authors' 308 QC-passing amplicons.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Human amplicons were cloned in the 3′ UTR of an EGFP reporter driven by the Hsp1a minimal promoter, using a modified STARR-seq orientation. The library was packaged in scAAV9(2YF), injected into the left prefrontal cortex of four P0 mouse brains, and collected at P7. Viral genomic DNA served as the delivery/input control and DNase-treated total-RNA cDNA provided the reporter output; Illumina amplicon sequencing was deduplicated and aligned to GRCh38. The paper also generated one matched high-cycle RNA technical replicate (L4 RNA, 35 PCR cycles), which is retained as a count column but excluded from the four-replicate activity summaries.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (78 of 78)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 78 definitions
amplicon_id
Unique numeric amplicon identifier from the authors' Amp_number field.
amplicon_name
Author-provided amplicon label, including number and locus/design token.
design_group
Library design group: GWAS, LD, FBDHS, or PutEnh.
chromosome
GRCh38 chromosome for the activity-table amplicon, with chr prefix.
start_grch38
GRCh38 start coordinate from the authors' activity table.
end_grch38
GRCh38 end coordinate from the authors' activity table.
length_bp
Inclusive coordinate span calculated as end_grch38 - start_grch38 + 1.
sequence_grch38
In-silico PCR amplicon sequence from supplementary file 9.
sequence_length_bp
Length in base pairs of sequence_grch38.
sequence_gc_fraction
GC fraction reported for the in-silico PCR sequence in supplementary file 9.
model_gc_fraction
Rounded GC covariate used in the authors' activity model.
previral_maxiprep_count
Raw count in the previral Maxiprep library.
previral_maxiprep_proportion
Previral Maxiprep count normalized to the total Maxiprep library count.
dna_rep_1_count
Raw viral genomic-DNA count for biological replicate L1.
dna_rep_2_count
Raw viral genomic-DNA count for biological replicate L2.
dna_rep_3_count
Raw viral genomic-DNA count for biological replicate L3.
dna_rep_4_count
Raw viral genomic-DNA count for biological replicate L4.
rna_rep_1_count
Raw reporter-RNA/cDNA count for biological replicate L1.
rna_rep_2_count
Raw reporter-RNA/cDNA count for biological replicate L2.
rna_rep_3_count
Raw reporter-RNA/cDNA count for biological replicate L3.
rna_rep_4_count
Raw reporter-RNA/cDNA count for biological replicate L4.
rna_rep_4_technical_35cycles_count
L4 reporter-RNA/cDNA count from the matched 35-PCR-cycle technical replicate; not used in summary activity.
dna_count_mean
Mean raw DNA count across L1-L4.
dna_count_sd
Standard deviation of raw DNA counts across L1-L4.
rna_count_mean
Mean raw RNA/cDNA count across L1-L4.
rna_count_sd
Standard deviation of raw RNA/cDNA counts across L1-L4.
dna_rep_1_proportion
L1 DNA count normalized to the total L1 DNA library count.
dna_rep_2_proportion
L2 DNA count normalized to the total L2 DNA library count.
dna_rep_3_proportion
L3 DNA count normalized to the total L3 DNA library count.
dna_rep_4_proportion
L4 DNA count normalized to the total L4 DNA library count.
rna_rep_1_proportion
L1 RNA/cDNA count normalized to the total L1 RNA library count.
rna_rep_2_proportion
L2 RNA/cDNA count normalized to the total L2 RNA library count.
rna_rep_3_proportion
L3 RNA/cDNA count normalized to the total L3 RNA library count.
rna_rep_4_proportion
L4 RNA/cDNA count normalized to the total L4 RNA library count.
dna_prop_mean
Mean normalized DNA proportion across L1-L4.
dna_prop_sd
Standard deviation of normalized DNA proportions across L1-L4.
rna_prop_mean
Mean normalized RNA proportion across L1-L4.
rna_prop_sd
Standard deviation of normalized RNA proportions across L1-L4.
activity_rep_1
L1 ratiometric activity, normalized RNA proportion divided by normalized DNA proportion.
activity_rep_2
L2 ratiometric activity, normalized RNA proportion divided by normalized DNA proportion.
activity_rep_3
L3 ratiometric activity, normalized RNA proportion divided by normalized DNA proportion.
activity_rep_4
L4 ratiometric activity, normalized RNA proportion divided by normalized DNA proportion.
mean_activity
Mean of L1-L4 ratiometric activities.
sd_activity
Standard deviation of L1-L4 ratiometric activities.
mean_ratio_of_means
RNA_prop_mean divided by DNA_prop_mean.
mean_ratio_replicate_activity
Mean ratiometric activity calculated from the four replicate activity ratios.
mean_ratio_sd
Standard deviation of the four replicate activity ratios.
activity_rank
Rank of mean_ratio_of_means among the QC-passing amplicons.
model_residual
Manual residual from the GC-adjusted linear activity model.
model_residual_z
Residual standardized to the fitted linear-model residual distribution.
model_p_value
One-tailed model p-value for increased RNA activity based on the standardized residual.
model_q_value_two_tailed
Empirical tail-area FDR q-value from the authors' residual analysis, reported as two-tailed.
model_fdr_lt_0_1
Author-provided boolean for the one-tailed model q-value meeting FDR<0.1.
wilcoxon_p_value
Two-tailed Wilcoxon rank-sum p-value comparing DNA and RNA proportions across replicates.
wilcoxon_conf_low
Lower confidence bound for the Wilcoxon activity comparison.
wilcoxon_conf_high
Upper confidence bound for the Wilcoxon activity comparison.
wilcoxon_bh_fdr
Benjamini-Hochberg-adjusted two-tailed Wilcoxon p-value.
wilcoxon_significance
Author descriptor of one-tailed Wilcoxon significance.
model_p_lt_0_05
Derived boolean indicating model_p_value<0.05 with positive model_residual.
ratiometric_active
Derived boolean indicating mean_ratio_replicate_activity>1.5 and mean_ratio_sd<mean_ratio_replicate_activity.
dnase_roadmap
Whether the amplicon intersects a Roadmap fetal-brain DNase hypersensitivity interval.
h3k27me3_roadmap
Whether the amplicon intersects a Roadmap H3K27me3 interval.
h3k36me3_roadmap
Whether the amplicon intersects a Roadmap H3K36me3 interval.
h3k4me1_roadmap
Whether the amplicon intersects a Roadmap H3K4me1 interval.
h3k4me3_roadmap
Whether the amplicon intersects a Roadmap H3K4me3 interval.
h3k9me3_roadmap
Whether the amplicon intersects a Roadmap H3K9me3 interval.
fetal_brain_tf_footprints_boca
Whether the amplicon intersects fetal-brain TF footprints from BOCA.
fetal_lung_tf_footprints_boca
Whether the amplicon intersects fetal-lung TF footprints from BOCA.
k562_tf_footprints_boca
Whether the amplicon intersects K562-cell TF footprints from BOCA.
atac_human_neurons_boca
Whether the amplicon intersects human-neuron ATAC-seq peaks from BOCA.
atac_human_glia_boca
Whether the amplicon intersects human-glia ATAC-seq peaks from BOCA.
conserved_element_boca
Whether the amplicon intersects a vertebrate conserved element annotation.
snp_count
Number of SNP observations associated with the amplicon in supplementary file 7.
snp_ids
Semicolon-separated unique SNP identifiers associated with the amplicon.
snp_positions_grch38
Semicolon-separated chr:position SNP coordinates on GRCh38.
allele_qc_snp_count
Number of associated SNP observations with Min_DNA_depth, Min_RNA_depth, and Minor_allele_freq_threshold all true.
allele_qc_snp_ids
Semicolon-separated SNP IDs meeting all three author allele-analysis QC flags.
qc_pass
Package QC retention flag; true for rows inherited from the authors' 308-row QC-filtered activity table.

Quality control

The authors first required at least 200 previral Maxiprep reads per amplicon, yielding 345 successfully cloned elements from the 408-element design. Downstream QC removed amplicons with fewer than 200 raw counts in any DNA sample or a mean DNA-library proportion below 2^-15 in any DNA sample, retaining 308/345 (89%) elements. The processed table uses that author-provided 308-row QC-filtered set, keeps four biological replicates, and excludes the L4 35-cycle technical RNA replicate from summary activity metrics. As validation, all 308 rows have unique IDs and finite source activity values; the packaged table reproduces the paper's 41 model p<0.05 calls, 17 model FDR<0.1 calls, 71 ratiometric calls, and 77 high-coverage/MAF-qualified SNP observations.

Curation notes

The full screen is the primary MPRA experiment; its human candidate sequences are assayed in mouse brain, so GRCh38 describes the library while the target organism and biosample describe the in vivo host. The allele-specific analysis is secondary and the paper notes limited power for moderate/small allelic effects; raw_data/supplementary_file_7_SNP_All_440.csv retains all 440 SNP observations, while the processed table aggregates the 413 observations associated with retained amplicons and the 77 observations meeting the paper's high-coverage/MAF flags. Seven retained amplicons had coordinate discrepancies between activity and in-silico PCR tables, so their unique primer-ID sequence records were used; two of those sequences are one base shorter than the activity coordinate span.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.