Experiment / E79Q0K9GFOther

Random 50-nt yeast 5′ UTR library HIS3 growth selection

Deep learning of the regulatory grammar of yeast 5′ untranslated regions from 500,000 random sequences

A pooled low-copy p415-CYC1-HIS3 plasmid library containing 50-nt random 5′ UTR inserts was transformed into BY4741 yeast lacking a native HIS3 copy. The abundance of 489,348 detected variants was measured before and after competitive histidine selection, providing a sequence-linked protein-expression proxy.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Input SD-Leu versus SD-His-Leu + 1.5 mM 3-AT competitive growth selection

The CYC1 promoter and terminator were held constant and LEU2 selected the plasmid. BY4741 transformants were competed for approximately 6.2 doublings in SD-His-Leu containing 1.5 mM 3-amino-1,2,4-triazole (3-AT); the paper reports three selection back-dilutions, while the GEO submission provides one aggregate t0/t1 count pair per sequence. The deposited growth_rate is numerically the natural-log selection/input enrichment after a pseudocount of one and library-size normalization; the processed table preserves that source value and adds a derived log2_enrichment column.

Massively parallel episomal plasmid growth-selection reporter: a low-copy p415-CYC1-HIS3 plasmid carried pooled 50-nt random 5′ UTR inserts immediately upstream of the HIS3 start codon. Variant abundance was quantified by sequencing plasmid DNA before and after growth in histidine-free medium, and selection/input enrichment was used as the reporter readout.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (10 of 10)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 10 definitions
element_id
Stable package identifier derived from the original GEO row index, formatted as random_NNNNNN.
source_index
Original unnamed row index in GSM2793752_Random_UTRs.csv.gz.
sequence
The tested 50-nt random 5′ UTR sequence.
sequence_length
Length of the tested UTR sequence in nucleotides.
input_count
Deposited t0 plasmid-DNA sequence count before histidine selection.
selected_count
Deposited t1 plasmid-DNA sequence count after histidine/3-AT selection.
input_fraction
Pseudocount-adjusted input_count divided by the total pseudocount-adjusted input count.
selected_fraction
Pseudocount-adjusted selected_count divided by the total pseudocount-adjusted selected count.
growth_rate_source
GEO-deposited growth_rate/enrichment value; numerically ln(selected_fraction/input_fraction) for this table.
log2_enrichment
Derived log2(selected_fraction/input_fraction), calculated from the raw counts with a pseudocount of one.

Quality control

The authors collapsed near-identical random sequences using a Hamming-distance threshold of less than 3, removed sequences shorter than 3 nt, aligned reads to the resulting synthetic reference, counted variant alignments, added a pseudocount of one, and normalized counts before calculating enrichment. The GEO table is already post-processing: all 489348 retained rows have unique 50-nt A/C/G/T sequences, nonnegative integer counts, and finite enrichment values, so no additional rows were removed. Zero selected counts are retained because they represent valid strong depletion measurements under the authors’ pseudocount procedure.

Curation notes

This is a pooled plasmid reporter selection and is included as an MPRA-like assay even though it does not use a conventional independent barcode RNA/DNA ratio. The paper and GEO describe the score as log2 enrichment, but the deposited growth_rate values match the natural logarithm of the pseudocount-normalized selection/input ratio; that unit discrepancy is preserved and made explicit rather than silently changing the source score. The 489348-row table is an aggregate submission without replicate-level counts or raw FASTQ files.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.