Experiment / E7YOBLYTYStandard STARR-seq

Anopheles coluzzii genome-wide STARR-seq enhancer screen in 4a3A cells

Comprehensive Genomic Discovery of Non-Coding Transcriptional Enhancers in the African Malaria Vector Anopheles coluzzii

A plasmid-based genome-wide STARR-seq screen tested randomly sheared genomic fragments from 60 wild Anopheles coluzzii collected in Burkina Faso in homologous 4a3A cells. Three independent biological replicates were analyzed, and the table contains the 3,288 merged enhancer peaks detected reproducibly across all three replicates.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Genomic DNA from 60 wild A. coluzzii was pooled, sheared to approximately 800 bp–1 kb, and cloned into the 3′ UTR of the pSTARR-seq_fly reporter downstream of a basal promoter. The library was transfected into A. coluzzii 4a3A cells; reporter cDNA and input plasmid DNA were sequenced on an Illumina HiSeq 2500 with 2 × 125-bp reads. Three biological replicates were independently transfected, harvested after 24 h, and analyzed by cDNA-versus-plasmid enrichment.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (38 of 38)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 38 definitions
enhancer_id
Stable package-local identifier assigned in original Supplementary Table S2 row order.
source_s2_row
1-based data-row number in Supplementary Table S2, excluding its three header rows.
chromosome_arm
Anopheles gambiae AgamP4 chromosome-arm label reported by the source.
merged_start
Start coordinate of the merged peak across the three biological replicates, using the source coordinate convention.
merged_end
End coordinate of the merged peak across the three biological replicates, using the source coordinate convention.
merged_width_bp
Source-reported merged peak width in base pairs; the source uses end - start + 1.
annotation
DataSheet 1 BED genomic feature annotation, such as intergenic, exon, UTR, or intron class.
upstream_gene
Gene reported by the source as upstream/nearest on the upstream side; NA indicates none reported.
distance_to_upstream_gene
Source-reported distance to the upstream gene; zero denotes overlap and direction/sign follows the source BED.
downstream_gene
Gene reported by the source as downstream/nearest on the downstream side; NA indicates none reported.
distance_to_downstream_gene
Source-reported distance to the downstream gene; zero denotes overlap and direction/sign follows the source BED.
rep1_start
Start coordinate of the replicate-1 peak, using the source coordinate convention.
rep1_end
End coordinate of the replicate-1 peak, using the source coordinate convention.
rep1_width_bp
Source-reported width of the replicate-1 peak.
rep1_cdna_coverage
Read coverage in the reporter cDNA library for biological replicate 1.
rep1_plasmid_coverage
Read coverage in the input plasmid DNA library for biological replicate 1.
rep1_p_value
Source peak-calling p-value for biological replicate 1.
rep1_fold_enrichment
Source cDNA/plasmid fold enrichment for biological replicate 1.
rep2_start
Start coordinate of the replicate-2 peak, using the source coordinate convention.
rep2_end
End coordinate of the replicate-2 peak, using the source coordinate convention.
rep2_width_bp
Source-reported width of the replicate-2 peak.
rep2_cdna_coverage
Read coverage in the reporter cDNA library for biological replicate 2.
rep2_plasmid_coverage
Read coverage in the input plasmid DNA library for biological replicate 2.
rep2_p_value
Source peak-calling p-value for biological replicate 2.
rep2_fold_enrichment
Source cDNA/plasmid fold enrichment for biological replicate 2.
rep3_start
Start coordinate of the replicate-3 peak, using the source coordinate convention.
rep3_end
End coordinate of the replicate-3 peak, using the source coordinate convention.
rep3_width_bp
Source-reported width of the replicate-3 peak.
rep3_cdna_coverage
Read coverage in the reporter cDNA library for biological replicate 3.
rep3_plasmid_coverage
Read coverage in the input plasmid DNA library for biological replicate 3.
rep3_p_value
Source peak-calling p-value for biological replicate 3.
rep3_fold_enrichment
Source cDNA/plasmid fold enrichment for biological replicate 3.
replicate_support
Number of biological replicates contributing to the merged peak; all rows equal 3.
mean_fold_enrichment
Arithmetic mean of the three source replicate fold-enrichment values; derived.
median_fold_enrichment
Median of the three source replicate fold-enrichment values; derived.
log2_mean_fold_enrichment
Base-2 logarithm of mean_fold_enrichment; derived activity summary.
mean_cdna_coverage
Arithmetic mean of the three replicate cDNA coverage values; derived.
mean_plasmid_coverage
Arithmetic mean of the three replicate plasmid-DNA coverage values; derived.

Quality control

The authors assessed sequencing quality with FastQC, mapped reads to the Anopheles gambiae AgamP4 assembly using BWA-MEM, retained properly mapped reads, removed supplementary alignments, and excluded Y, UNKN, and mitochondrial chromosomes from peak calling. BasicSTARRseq getPeaks used minQuantile=0.99, peakWidth=500, maxPval=0.001, model=2, and the authors retained peaks with cDNA/plasmid fold enrichment >=3. Merged peaks were required to overlap by at least 250 bp across all three biological replicates. The packaged table retains all 3,288 peaks passing these author-defined filters; no additional rows were removed beyond validation of unique S2/BED coordinate joins and numeric fields.

Curation notes

This package represents one MPRA experiment. The manual luciferase validation panel reported in the paper was not included as a second experiment because it is a low-throughput follow-up assay rather than an MPRA. ENA read-run metadata is included, but raw FASTQ files are intentionally omitted. The S2/BED coordinate join was complete and unique for all 3,288 rows. The supplied BED has one malformed duplicate attribute at 2L:48474399-48475406 (two EnrichmentRep3 values); the first agrees with Supplementary Table S2, which was used for processed replicate activity. The paper describes the 4a3A line as molecularly typed A. coluzzii, while its Cellosaurus record is indexed under the legacy A. gambiae species of origin.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.