Experiment / E5L5AXHRPSort-Seq / Flow-Seq MPRA

Combinatorial 26-element regulatory cassette library: FACS/Nanopore sort-seq

Permutational analysis of Saccharomyces cerevisiae regulatory elements

A directed Golden Gate library of approximately 400,000 episomal yeast reporter plasmids randomly combined 26 endogenous UAS enhancers, 25 cloned core promoters, 26 5′ UTRs and 26 3′ UTR/terminators upstream of mRuby2. W-303 yeast cells were sorted into no-, low-, medium- and high-mRuby2 expression bins, and the cassette identities were recovered by Oxford Nanopore sequencing.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Glucose-rich YMD medium

The reporter library was an episomal ARS/CEN/URA3-derived plasmid system assembled from modular parts and expressed mRuby2; constitutive mTagEBFP2-2 under the RPL18b promoter was used for FACS gating. Four sorted pools were sequenced with Oxford Nanopore barcodes NB01–NB04. The readout is a FACS-bin distribution for reporter protein expression rather than a DNA/RNA abundance ratio. The processed table reports all six pairwise regulatory-element classes by pooling the other two complete fragments.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (28 of 28)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 28 definitions
pair_type
Pairwise regulatory-element class; the other two complete fragments were pooled.
element_1_type
Controlled type of the first element in the pair.
element_1
Gene-derived regulatory-fragment name for the first element in the pair.
element_2_type
Controlled type of the second element in the pair.
element_2
Gene-derived regulatory-fragment name for the second element in the pair.
presort_read_count
Complete four-fragment cassette assignments in the unsorted library sample.
no_expression_read_count
Complete cassette assignments from NB01, the no-mRuby2-expression FACS pool.
low_expression_read_count
Complete cassette assignments from NB02, the low-mRuby2-expression FACS pool.
medium_expression_read_count
Complete cassette assignments from NB03, the medium-mRuby2-expression FACS pool.
high_expression_read_count
Complete cassette assignments from NB04, the high-mRuby2-expression FACS pool.
sorted_evidence_reads
Sum of the four sorted-pool read counts used for the QC cutoff.
n_unique_complete_cassettes
Number of distinct complete four-fragment cassettes contributing to this pair.
raw_no_expression_fraction
NB01 divided by sorted_evidence_reads.
raw_low_expression_fraction
NB02 divided by sorted_evidence_reads.
raw_medium_expression_fraction
NB03 divided by sorted_evidence_reads.
raw_high_expression_fraction
NB04 divided by sorted_evidence_reads.
facs_adjusted_no_expression_fraction
Cell-count-adjusted NB01 fraction after scaling each pool by its number of sorted cells and renormalizing.
facs_adjusted_low_expression_fraction
Cell-count-adjusted NB02 fraction after scaling each pool by its number of sorted cells and renormalizing.
facs_adjusted_medium_expression_fraction
Cell-count-adjusted NB03 fraction after scaling each pool by its number of sorted cells and renormalizing.
facs_adjusted_high_expression_fraction
Cell-count-adjusted NB04 fraction after scaling each pool by its number of sorted cells and renormalizing.
raw_weighted_expression_score
Weighted sorted-bin score using bin values [1, 60, 120, 250], divided by 250; unadjusted fraction scale.
facs_adjusted_weighted_expression_score
Weighted score using the cell-count-adjusted bin fractions and bin values [1, 60, 120, 250], divided by 250.
inferred_mean_expression_au
Mean expression estimate from a deterministic fit of the public notebook's fixed-log-SD 0.95 log-normal model to the adjusted bin fractions.
inferred_mean_expression_2d_au
Mean expression selected from the public notebook's coarse two-dimensional grid fit.
inferred_log_sd_2d
Log-space standard deviation selected by the coarse two-dimensional grid fit.
model_fit_sse_fixed_log_sd
Sum of squared errors for the fixed-log-SD 0.95 fit to adjusted bin fractions.
model_fit_sse_2d
Sum of squared errors for the coarse two-dimensional mean/log-SD fit.
qc_pass
TRUE for a row retained after the complete-fragment and sorted-evidence filters.

Quality control

The authors designed FACS gates using four control strains and collected 24,540,211 cells; the four bins contained 15,663,845, 6,462,912, 1,886,385 and 527,069 cells, respectively. The public analysis grouped only reads with all four nonzero fragment assignments, then used an evidence cutoff of 100 sorted reads for pairwise activity displays; this package retains rows with sorted_evidence_reads >= 100. Zero-coded or otherwise incomplete fragment assignments were excluded before aggregation. The resulting table contains 3,113 QC-passing pair rows: 521 UAS–CORE, 492 UAS–UTR5, 444 UAS–UTR3, 599 CORE–UTR5, 545 CORE–UTR3 and 512 UTR5–UTR3 rows. Complete-fragment totals retained from the raw files are 1,359,713 presort reads and 739,101, 241,786, 67,059 and 169,098 reads in NB01–NB04; 109,918 unique complete cassettes were observed across these files.

Curation notes

The assayed cells were the W-303 derivative ROY5634, while all regulatory fragments were PCR amplified from S288c genomic DNA; the reference assembly was not stated. Fragment indexes are mapped using the public repository's 1-based list. The paper reports that the HIS3 core promoter could not be cloned and that ADH1 and PHO5 enhancer fragments were missing; the raw assignment files contain a small presort-only PHO5 UAS signal but no corresponding sorted-pool reads, and no ADH1 UAS assignments. Because this is a sort-seq reporter assay, the activity scores should be interpreted as protein-expression-bin estimates, not RNA/DNA log2 ratios.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.