Experiment / E23W0ZTH9Targeted Genomic Integration MPRA

duf:CD2+ / GFP+ muscle founder-cell eFS

Highly parallel assays of tissue-specific enhancers in whole Drosophila embryos

The second injection batch of the cCRM library was assayed in stage 11–12 embryos carrying the duf:CD2 muscle-founder marker, and the rare GFP-positive CD2-positive population was pooled before sequencing. The pooled target was compared with five size-matched duf:CD2 input-control samples.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Each approximately 1-kb cCRM was cloned upstream of a nuclear EGFP reporter driven by an Hsp70 minimal promoter in pEFS-Dest and integrated one per transformant line by phiC31 into the attP40 landing site. Stage 11–12 embryo cells were stained with Alexa647 anti-rat CD2, separated by CD2/GFP FACS gates, and the cCRM inserts were recovered from genomic DNA by nested PCR and sequenced as 50-base single-end Illumina HiSeq 2000 reads mapped with segemehl. This is a region-focused genomic integration assay with no molecular barcodes or allele pairs. For this contrast, the target samples were duf:CD2+/GFP+ pooled and the input reference samples were duf:CD2 input control rep. 1–5.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (33 of 33)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 33 definitions
element_id
Unique cCRM/library element identifier from the authors' Supplementary Tables 1 and 4.
chromosome
Drosophila chromosome reported for the cCRM in the dm3 coordinate string.
start
One-based genomic start coordinate reported by the authors.
end
Genomic end coordinate reported by the authors.
coordinate_span_bp
Coordinate-derived span in base pairs, calculated from the reported start and end.
sequence
Candidate cis-regulatory module sequence supplied in Supplementary Table 1.
sequence_length
Number of bases in the supplied cCRM sequence.
gc_fraction
Fraction of sequence bases that are G or C.
forward_primer
Forward primer sequence supplied for recovery/amplification of the cCRM.
reverse_primer
Reverse primer sequence supplied for recovery/amplification of the cCRM.
library_source_category
Library/source category reported for the cCRM.
input_count_rep1
Raw overlap-weighted mapped cCRM Count from input sample replicate 1; the exact sample title is documented in agent_comments and raw_data/E-GEOD-41503.sdrf.txt.
input_count_rep2
Raw overlap-weighted mapped cCRM Count from input sample replicate 2; the exact sample title is documented in agent_comments and raw_data/E-GEOD-41503.sdrf.txt.
input_count_rep3
Raw overlap-weighted mapped cCRM Count from input sample replicate 3; the exact sample title is documented in agent_comments and raw_data/E-GEOD-41503.sdrf.txt.
input_count_rep4
Raw overlap-weighted mapped cCRM Count from input sample replicate 4; the exact sample title is documented in agent_comments and raw_data/E-GEOD-41503.sdrf.txt.
input_count_rep5
Raw overlap-weighted mapped cCRM Count from input sample replicate 5; the exact sample title is documented in agent_comments and raw_data/E-GEOD-41503.sdrf.txt.
target_count_rep1
Raw overlap-weighted mapped cCRM Count from target sample replicate 1; the exact sample title is documented in agent_comments and raw_data/E-GEOD-41503.sdrf.txt.
input_count_mean
Arithmetic mean of the raw overlap-weighted mapped cCRM Count values across input replicates.
target_count_mean
Arithmetic mean of the raw overlap-weighted mapped cCRM Count values across target replicates.
input_cpm_mean
Mean input abundance after normalizing each input file's cCRM Count values to counts per million.
target_cpm_mean
Mean target abundance after normalizing each target file's cCRM Count values to counts per million.
normalized_target_input_ratio
Recomputed ratio of mean target CPM to mean input CPM; this is not the authors' DESeq statistic.
normalized_log2_enrichment
Base-2 logarithm of normalized_target_input_ratio; zero target is represented as -Inf and zero input as Inf.
input_detected_replicates
Number of input replicates meeting the paper's detection thresholds: both ends detected, at least 5 center positions covered, and at least 10 total reads.
target_detected_replicates
Number of target replicates meeting the paper's detection thresholds: both ends detected, at least 5 center positions covered, and at least 10 total reads.
author_input_abundance
Input abundance reported by the authors in Supplementary Table 4 for this contrast.
author_log2_fold_change
Authors' DESeq log2 fold change for target versus input; retained as reported.
author_padj
Authors' DESeq multiple-testing-adjusted P value for the target-versus-input contrast.
author_activity_call
Packaging call: active when author_padj is below 0.1, otherwise not_significant; blank when no author adjusted P value is available.
redfly_overlap_bp
Number of base pairs overlapping a RedFly annotated regulatory region, as reported by the authors.
hot_region_overlap
Whether the cCRM overlaps a reported hot region, as reported by the authors.
validation_call
Summary validation call reported by the authors.
validation_finding
Validation result or explanatory finding reported by the authors.

Quality control

The authors required at least 1 read from each cCRM end, at least 5 center positions covered, and at least 10 total reads; matched random genomic windows gave a cCRM-detection FDR below 5×10−5. cCRMs not detected in any input sample replicate were filtered before DESeq analysis, and the authors used DESeq size-factor estimation with adjusted P <0.1 as the activity threshold. For packaging, rows were retained when the corresponding Supplementary Table 4 inputAbundance was numeric and the Supplementary Table 1 record had a valid sequence and coordinate. Across the deposited samples, 92.50–98.59% of filtered reads were assigned to cCRMs. Inactive and depleted elements were retained; zero-target results are represented as -Inf.

Curation notes

This is one of six child contrasts from the same eFS study/library, not an allele-specific variant experiment. Target sample titles: duf:CD2+/GFP+ pooled; input sample titles: duf:CD2 input control rep. 1–5. The raw title-to-file and GEO/BioStudies sample metadata are retained in raw_data/E-GEOD-41503.sdrf.txt, and the mapped cCRM Count files are in raw_data/geo_processed/. The normalized_* columns are recomputed from each raw file's Count divided by that file's total mapped cCRM Count and are not replacements for the authors' DESeq statistics; use author_log2_fold_change and author_padj for the published contrast. Source NA/None annotations are empty CSV fields, while author -Inf/Inf values are retained. The source coordinate span and sequence length differ by one base for some cCRMs; both supplied sequence and coordinate values are preserved. The single target file represents pooled rare GFP+CD2+ muscle-founder-cell collections; the five input samples are size-matched duf:CD2 controls.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.