Large-scale RPL8A 5′-UTR mutant library (2,041 variants)
Deciphering the rules by which 5′-UTR sequences affect protein expression in yeastA genomically integrated RPL8A promoter–YFP reporter library in Saccharomyces cerevisiae varied the 10 nucleotides immediately upstream of the YFP start codon. Cells were sorted into 24 YFP/mCherry fluorescence bins, and bin-specific sequencing was used to estimate mean protein abundance for each variant.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
Randomized 5′-UTR sequence at positions −10 to −1; basal growth conditions
The library was integrated into yeast as an RPL8A promoter–YFP reporter and paired with an identical TEF2 promoter–mCherry control cassette. FACS gates separated cells into 24 bins by YFP/mCherry ratio; each bin was amplified with a unique 5-bp barcode and pooled for SOLiD sequencing. Mean protein abundance was calculated as the normalized, bin-fluorescence-weighted mean for each 5′-UTR variant.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (14 of 14)
| Row | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ||||||||||||||
| 2 | ||||||||||||||
| 3 | ||||||||||||||
| 4 | ||||||||||||||
| 5 | ||||||||||||||
| 6 | ||||||||||||||
| 7 | ||||||||||||||
| 8 | ||||||||||||||
| 9 | ||||||||||||||
| 10 | ||||||||||||||
| 11 | ||||||||||||||
| 12 | ||||||||||||||
| 13 | ||||||||||||||
| 14 | ||||||||||||||
| 15 | ||||||||||||||
| 16 | ||||||||||||||
| 17 | ||||||||||||||
| 18 | ||||||||||||||
| 19 | ||||||||||||||
| 20 | ||||||||||||||
| 21 | ||||||||||||||
| 22 | ||||||||||||||
| 23 | ||||||||||||||
| 24 | ||||||||||||||
| 25 | ||||||||||||||
| 26 | ||||||||||||||
| 27 | ||||||||||||||
| 28 | ||||||||||||||
| 29 | ||||||||||||||
| 30 | ||||||||||||||
| 31 | ||||||||||||||
| 32 | ||||||||||||||
| 33 | ||||||||||||||
| 34 | ||||||||||||||
| 35 | ||||||||||||||
| 36 | ||||||||||||||
| 37 | ||||||||||||||
| 38 | ||||||||||||||
| 39 | ||||||||||||||
| 40 | ||||||||||||||
| 41 | ||||||||||||||
| 42 | ||||||||||||||
| 43 | ||||||||||||||
| 44 | ||||||||||||||
| 45 | ||||||||||||||
| 46 | ||||||||||||||
| 47 | ||||||||||||||
| 48 | ||||||||||||||
| 49 | ||||||||||||||
| 50 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 14 definitions
- variant_id
- Stable row identifier assigned in Dataset S1 order
- sequence_variant_10mer
- The tested 10-nt sequence replacing RPL8A 5′-UTR positions −10 to −1
- full_5utr_sequence
- The 17-nt reporter 5′-UTR sequence, consisting of constant AAAACAA followed by the tested 10-mer
- is_reference_sequence
- Whether the tested 10-mer matches the reported native RPL8A sequence CTAATTCGAA
- substitutions_vs_reference
- Hamming distance from native RPL8A 10-mer CTAATTCGAA
- gc_fraction
- Fraction of bases in the tested 10-mer that are G or C
- kozak_minus3_minus1
- Three nucleotides immediately upstream of the reporter start codon (positions −3 to −1)
- upstream_aug_count
- Count of ATG triplets in the 17-nt upstream sequence
- in_frame_upstream_aug_count
- Count of upstream ATGs aligned in-frame with the main reporter ORF
- out_of_frame_upstream_aug_count
- Count of upstream ATGs out-of-frame with the main reporter ORF
- mean_protein_abundance
- Study-reported mean YFP/mCherry protein-abundance estimate from the 24-bin sequencing assay
- total_reads
- Total mapped sequencing reads for the variant in Dataset S1
- relative_protein_abundance_vs_reference
- Mean protein abundance divided by the native CTAATTCGAA reference abundance
- log2_fold_change_vs_reference
- Log2 of relative protein abundance versus the native CTAATTCGAA reference
Quality control
Dataset S1 entries were retained only when the 5′-UTR variable sequence was unique, exactly 10 nt long, contained only A/C/G/T, had finite positive mean protein abundance, and had at least 100 mapped reads, matching the paper's stated minimum coverage. All 2,041 deposited variants passed these checks, so no rows were excluded. The study's own processing normalized bin and variant counts and removed small-count tails representing at most 5% of each variant's reads before calculating abundance; the mean estimates were also validated against 84 individually isolated strains (Pearson R² = 0.98).
Curation notes
The study uses a genomically integrated reporter and a FACS/bin-specific sequencing readout, so it is classified as Sort-Seq / Flow-Seq MPRA rather than episomal plasmid MPRA. The processed table represents the primary 2,041-variant pooled assay deposited as Dataset S1. Two follow-up pools (44 A/C-rich variants and 65 G/C-rich variants) were measured individually by Sanger sequencing and microplate reader and were not treated as separate MPRA experiments. The reported native RPL8A 5′-UTR is AAAACAACTAATTCGAA, giving CTAATTCGAA as the variable native 10-mer. No biological replicate-level counts, bin matrix, or p-values were deposited with Dataset S1, so those fields are not inferred.