Experiment / E9OHYUNS7Sort-Seq / Flow-Seq MPRA

Designed NBT yeast promoter MPRA

A community effort to optimize sequence-based deep learning models of gene regulation

A separately measured designed promoter library of approximately 71,000 constructs was assayed in S288C ΔURA3 yeast using the same dual-reporter FACS-bin workflow as the random library. The library contains native yeast fragments, random controls, high- and low-expression designs, challenging sequences, SNV perturbations, and motif perturbation or tiling constructs.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Chardonnay grape must

Designed NBT oligos were cloned into the same yeast dual reporter with an 80-bp variable promoter region upstream of YFP and constitutive pTEF2-RFP for extrinsic-noise control. The library was transformed at approximately 10^5 E. coli transformants, providing more than 10× library coverage, grown in Chardonnay grape must, sorted by log2(RFP/YFP) into 18 uniform bins, sequenced with 2 × 76-bp paired-end reads, and quantified with MAUDE from the bin abundances and estimated initial sequence abundances.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (17 of 17)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 17 definitions
element_id
Package-generated unique identifier for the scored NBT construct, assigned in source-row order.
source_row_index
Zero-based row index in the author-released filtered_test_data_with_MAUDE_expression table.
sequence
Full 110-nt observed promoter reporter sequence, including fixed distal and proximal scaffold flanks and the designed 80-bp variable region.
variable_sequence
The 80-bp designed sequence between the fixed reporter scaffold flanks.
sequence_length
Length of the full reporter promoter sequence in nucleotides; all retained rows are 110.
variable_length
Length of the extracted designed sequence between fixed flanks; all retained rows are 80.
mean_expression
Author-provided MAUDE expression score for the designed construct; larger values indicate higher reporter expression.
gc_fraction
Fraction of bases in the full reporter sequence that are G or C.
variable_gc_fraction
Fraction of bases in the designed variable sequence that are G or C.
test_subsets
Named test-library subset annotation joined from the official archive; semicolon-separated when a sequence belongs to multiple subsets, or Uncategorized held-out when no named archive file contains it.
source_tags
Author-designed construct or pair tags from the subset archive; semicolon-separated when multiple tags map to one sequence.
pair_ids
Semicolon-separated identifiers for SNV, motif-perturbation, or motif-tiling pair records involving this sequence.
pair_roles
Semicolon-separated roles corresponding positionally to pair_ids: alternate or reference.
paired_element_ids
Semicolon-separated package element IDs of the paired construct, positionally aligned to pair_ids.
paired_mean_expressions
Semicolon-separated author expression scores for the paired constructs, positionally aligned to pair_ids.
expression_deltas_vs_paired
Current construct mean_expression minus its paired construct expression, positionally aligned to pair_ids.
edit_distances_to_paired
Number of differing bases in the 80-bp variable sequence versus each paired construct, positionally aligned to pair_ids.

Quality control

The authors aligned designed-library reads directly to the ordered promoter sequences and used MAUDE to estimate expression from sorting-bin abundance, with initial abundance estimated from the average relative abundance across bins. Package QC retained rows only when the source had two tab-separated fields, a finite score, an exactly 110-nt A/C/G/T-only sequence with the expected reporter flanks, and no duplicate sequence. All 71,103 source rows passed. Test-subset and allele-pair annotations were joined from the official Zenodo test_subset_ids archive; rows not present in one of its named subset files were explicitly labeled Uncategorized held-out rather than discarded.

Curation notes

The processed table contains all 71,103 scored NBT constructs and 52,416 pair records (104,832 directional sequence-to-pair links) derived from the official SNV, motif-perturbation, and motif-tiling annotation files. Pair deltas are calculated from the author MAUDE score as current minus paired score; they are not re-estimated from raw reads. The test-subset archive does not annotate 8,911 scored rows, which are retained and labeled Uncategorized held-out. This is a diverse designed model-evaluation library rather than a single genomic locus or a pure disease-variant library, so reference_genome is reported from GEO as R64 while region_of_interest is null.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.