Experiment / E8NQ46DKXSort-Seq / Flow-Seq MPRA

Random N80 yeast promoter library MPRA

A community effort to optimize sequence-based deep learning models of gene regulation

A high-complexity library of nominally 80-bp random DNA inserts was placed between fixed distal and proximal promoter scaffold sequences upstream of YFP in a dual reporter and assayed in S288C ΔURA3 yeast. The package combines the author-released Zenodo train and validation sequence-expression tables for the random N80 library.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Chardonnay grape must

Random DNA was cloned into the yeast_Dual-Reporter plasmid (Addgene 127546) between a distal scaffold (pT) and proximal scaffold (pA), upstream of YFP. Constitutive RFP under pTEF2 was used to control extrinsic noise; S288C ΔURA3 yeast were grown in Chardonnay grape must, sorted by log2(RFP/YFP) on a Beckman-Coulter MoFlo Astrios into 18 uniform bins in three six-bin batches, and promoter inserts were sequenced with 2 × 76-bp paired-end reads.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (10 of 10)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 10 definitions
element_id
Package-generated unique identifier for the retained N80 promoter construct.
source_split
Official Zenodo split from which the row was taken: train or validation.
source_row_index
Zero-based row index within the corresponding official source split before package QC.
sequence
Observed full reporter promoter sequence, including the constant distal and proximal scaffold flanks and the random insert.
variable_sequence
Sequence between the fixed reporter scaffold flanks; nominally the 80-bp randomized insert, with observed length variation preserved.
sequence_length
Length of the observed full promoter sequence in nucleotides.
variable_length
Length of the extracted sequence between the fixed scaffold flanks in nucleotides.
mean_expression
Author-provided mean expression score calculated as the read-count-weighted mean across 18 FACS bins; larger values indicate higher YFP expression.
gc_fraction
Fraction of bases in the full observed promoter sequence that are G or C.
variable_gc_fraction
Fraction of bases in the extracted variable sequence that are G or C.

Quality control

Author processing aligned paired reads using a 40 ± 15-bp central overlap, discarded reads that failed the overlap constraint, clustered observed N80 promoters with Bowtie2, selected the highest-read representative for each cluster, and calculated expression as the read-count-weighted mean of the 18 sorting bins. Package QC required exactly two tab-separated fields, a finite expression value, the expected distal and proximal reporter scaffold flanks, and an A/C/G/T-only sequence. Of 6,739,250 source rows (6,065,325 train and 673,925 validation), 6,715,190 passed (6,043,623 train and 671,567 validation); 24,060 rows containing ambiguous bases were excluded. Observed sequence lengths of 83–142 nt were retained because the author-collapsed release contains length variation; all retained sequences were unique.

Curation notes

This is the study's primary random synthetic promoter reporter experiment, not a GWAS or allele-pair library. The paper reports 6,739,258 sequence-expression pairs, while the current Zenodo train.txt plus val.txt release contains 6,739,250 rows; no rows were added to reconcile that source-release discrepancy. The final table contains the clean A/C/G/T subset of the official release and preserves the author score and train/validation provenance. The raw sequencing reads are intentionally not included.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.