Experiment / E8G9WT411Sort-Seq / Flow-Seq MPRA

Large-scale RPL8A 5′-UTR mutant library (2,041 variants)

Deciphering the rules by which 5′-UTR sequences affect protein expression in yeast

A genomically integrated RPL8A promoter–YFP reporter library in Saccharomyces cerevisiae varied the 10 nucleotides immediately upstream of the YFP start codon. Cells were sorted into 24 YFP/mCherry fluorescence bins, and bin-specific sequencing was used to estimate mean protein abundance for each variant.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Randomized 5′-UTR sequence at positions −10 to −1; basal growth conditions

The library was integrated into yeast as an RPL8A promoter–YFP reporter and paired with an identical TEF2 promoter–mCherry control cassette. FACS gates separated cells into 24 bins by YFP/mCherry ratio; each bin was amplified with a unique 5-bp barcode and pooled for SOLiD sequencing. Mean protein abundance was calculated as the normalized, bin-fluorescence-weighted mean for each 5′-UTR variant.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (14 of 14)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 14 definitions
variant_id
Stable row identifier assigned in Dataset S1 order
sequence_variant_10mer
The tested 10-nt sequence replacing RPL8A 5′-UTR positions −10 to −1
full_5utr_sequence
The 17-nt reporter 5′-UTR sequence, consisting of constant AAAACAA followed by the tested 10-mer
is_reference_sequence
Whether the tested 10-mer matches the reported native RPL8A sequence CTAATTCGAA
substitutions_vs_reference
Hamming distance from native RPL8A 10-mer CTAATTCGAA
gc_fraction
Fraction of bases in the tested 10-mer that are G or C
kozak_minus3_minus1
Three nucleotides immediately upstream of the reporter start codon (positions −3 to −1)
upstream_aug_count
Count of ATG triplets in the 17-nt upstream sequence
in_frame_upstream_aug_count
Count of upstream ATGs aligned in-frame with the main reporter ORF
out_of_frame_upstream_aug_count
Count of upstream ATGs out-of-frame with the main reporter ORF
mean_protein_abundance
Study-reported mean YFP/mCherry protein-abundance estimate from the 24-bin sequencing assay
total_reads
Total mapped sequencing reads for the variant in Dataset S1
relative_protein_abundance_vs_reference
Mean protein abundance divided by the native CTAATTCGAA reference abundance
log2_fold_change_vs_reference
Log2 of relative protein abundance versus the native CTAATTCGAA reference

Quality control

Dataset S1 entries were retained only when the 5′-UTR variable sequence was unique, exactly 10 nt long, contained only A/C/G/T, had finite positive mean protein abundance, and had at least 100 mapped reads, matching the paper's stated minimum coverage. All 2,041 deposited variants passed these checks, so no rows were excluded. The study's own processing normalized bin and variant counts and removed small-count tails representing at most 5% of each variant's reads before calculating abundance; the mean estimates were also validated against 84 individually isolated strains (Pearson R² = 0.98).

Curation notes

The study uses a genomically integrated reporter and a FACS/bin-specific sequencing readout, so it is classified as Sort-Seq / Flow-Seq MPRA rather than episomal plasmid MPRA. The processed table represents the primary 2,041-variant pooled assay deposited as Dataset S1. Two follow-up pools (44 A/C-rich variants and 65 G/C-rich variants) were measured individually by Sanger sequencing and microplate reader and were not treated as separate MPRA experiments. The reported native RPL8A 5′-UTR is AAAACAACTAATTCGAA, giving CTAATTCGAA as the variable native 10-mer. No biological replicate-level counts, bin matrix, or p-values were deposited with Dataset S1, so those fields are not inferred.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.