Random 50-nt yeast 5′ UTR library HIS3 growth selection
Deep learning of the regulatory grammar of yeast 5′ untranslated regions from 500,000 random sequencesA pooled low-copy p415-CYC1-HIS3 plasmid library containing 50-nt random 5′ UTR inserts was transformed into BY4741 yeast lacking a native HIS3 copy. The abundance of 489,348 detected variants was measured before and after competitive histidine selection, providing a sequence-linked protein-expression proxy.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
Input SD-Leu versus SD-His-Leu + 1.5 mM 3-AT competitive growth selection
The CYC1 promoter and terminator were held constant and LEU2 selected the plasmid. BY4741 transformants were competed for approximately 6.2 doublings in SD-His-Leu containing 1.5 mM 3-amino-1,2,4-triazole (3-AT); the paper reports three selection back-dilutions, while the GEO submission provides one aggregate t0/t1 count pair per sequence. The deposited growth_rate is numerically the natural-log selection/input enrichment after a pseudocount of one and library-size normalization; the processed table preserves that source value and adds a derived log2_enrichment column.
Massively parallel episomal plasmid growth-selection reporter: a low-copy p415-CYC1-HIS3 plasmid carried pooled 50-nt random 5′ UTR inserts immediately upstream of the HIS3 start codon. Variant abundance was quantified by sequencing plasmid DNA before and after growth in histidine-free medium, and selection/input enrichment was used as the reporter readout.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (10 of 10)
| Row | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ||||||||||
| 2 | ||||||||||
| 3 | ||||||||||
| 4 | ||||||||||
| 5 | ||||||||||
| 6 | ||||||||||
| 7 | ||||||||||
| 8 | ||||||||||
| 9 | ||||||||||
| 10 | ||||||||||
| 11 | ||||||||||
| 12 | ||||||||||
| 13 | ||||||||||
| 14 | ||||||||||
| 15 | ||||||||||
| 16 | ||||||||||
| 17 | ||||||||||
| 18 | ||||||||||
| 19 | ||||||||||
| 20 | ||||||||||
| 21 | ||||||||||
| 22 | ||||||||||
| 23 | ||||||||||
| 24 | ||||||||||
| 25 | ||||||||||
| 26 | ||||||||||
| 27 | ||||||||||
| 28 | ||||||||||
| 29 | ||||||||||
| 30 | ||||||||||
| 31 | ||||||||||
| 32 | ||||||||||
| 33 | ||||||||||
| 34 | ||||||||||
| 35 | ||||||||||
| 36 | ||||||||||
| 37 | ||||||||||
| 38 | ||||||||||
| 39 | ||||||||||
| 40 | ||||||||||
| 41 | ||||||||||
| 42 | ||||||||||
| 43 | ||||||||||
| 44 | ||||||||||
| 45 | ||||||||||
| 46 | ||||||||||
| 47 | ||||||||||
| 48 | ||||||||||
| 49 | ||||||||||
| 50 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 10 definitions
- element_id
- Stable package identifier derived from the original GEO row index, formatted as random_NNNNNN.
- source_index
- Original unnamed row index in GSM2793752_Random_UTRs.csv.gz.
- sequence
- The tested 50-nt random 5′ UTR sequence.
- sequence_length
- Length of the tested UTR sequence in nucleotides.
- input_count
- Deposited t0 plasmid-DNA sequence count before histidine selection.
- selected_count
- Deposited t1 plasmid-DNA sequence count after histidine/3-AT selection.
- input_fraction
- Pseudocount-adjusted input_count divided by the total pseudocount-adjusted input count.
- selected_fraction
- Pseudocount-adjusted selected_count divided by the total pseudocount-adjusted selected count.
- growth_rate_source
- GEO-deposited growth_rate/enrichment value; numerically ln(selected_fraction/input_fraction) for this table.
- log2_enrichment
- Derived log2(selected_fraction/input_fraction), calculated from the raw counts with a pseudocount of one.
Quality control
The authors collapsed near-identical random sequences using a Hamming-distance threshold of less than 3, removed sequences shorter than 3 nt, aligned reads to the resulting synthetic reference, counted variant alignments, added a pseudocount of one, and normalized counts before calculating enrichment. The GEO table is already post-processing: all 489348 retained rows have unique 50-nt A/C/G/T sequences, nonnegative integer counts, and finite enrichment values, so no additional rows were removed. Zero selected counts are retained because they represent valid strong depletion measurements under the authors’ pseudocount procedure.
Curation notes
This is a pooled plasmid reporter selection and is included as an MPRA-like assay even though it does not use a conventional independent barcode RNA/DNA ratio. The paper and GEO describe the score as log2 enrichment, but the deposited growth_rate values match the natural logarithm of the pseudocount-normalized selection/input ratio; that unit discrepancy is preserved and made explicit rather than silently changing the source score. The 489348-row table is an aggregate submission without replicate-level counts or raw FASTQ files.