Experiment / E8GFISDWEOther

Native yeast 5′ UTR fragment HIS3 growth selection

Deep learning of the regulatory grammar of yeast 5′ untranslated regions from 500,000 random sequences

A pooled low-copy p415-CYC1-HIS3 plasmid library containing up to 50-nt fragments from known Saccharomyces cerevisiae native 5′ UTRs was transformed into BY4741 yeast lacking a native HIS3 copy. The abundance of 11,856 detected native fragments was measured before and after competitive histidine selection to estimate sequence-dependent protein expression.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Input SD-Leu versus SD-His-Leu + 1.5 mM 3-AT competitive growth selection

The CYC1 promoter and terminator were held constant and LEU2 selected the plasmid. The library comprised 50-nt native 5′ UTR segments with 25-nt overlap for UTRs longer than 50 nt and shorter fragments for UTRs shorter than 50 nt. BY4741 transformants were competed for approximately 6.2 doublings in SD-His-Leu containing 1.5 mM 3-amino-1,2,4-triazole (3-AT); the paper reports three selection back-dilutions, while the GEO submission provides one aggregate t0/t1 count pair per fragment. The processed table preserves the GEO source score and adds pseudocount-normalized fractions and a derived log2 enrichment.

Massively parallel episomal plasmid growth-selection reporter: a low-copy p415-CYC1-HIS3 plasmid carried pooled native 5′ UTR fragments immediately upstream of the HIS3 start codon. Fragment abundance was quantified by sequencing plasmid DNA before and after growth in histidine-free medium, and selection/input enrichment was used as the reporter readout.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (11 of 11)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 11 definitions
element_id
Stable package identifier derived from the original GEO row index, formatted as native_NNNNNN.
source_index
Original unnamed row index in GSM2793754_Native_UTRs.csv.gz.
utr_name
Source native UTR identifier, typically encoding the yeast gene and fragment coordinates.
sequence
The tested native Saccharomyces cerevisiae 5′ UTR fragment.
sequence_length
Length of the tested UTR fragment in nucleotides.
input_count
Deposited t0 plasmid-DNA sequence count before histidine selection.
selected_count
Deposited t1 plasmid-DNA sequence count after histidine/3-AT selection.
input_fraction
Pseudocount-adjusted input_count divided by the total pseudocount-adjusted input count.
selected_fraction
Pseudocount-adjusted selected_count divided by the total pseudocount-adjusted selected count.
growth_rate_source
GEO-deposited growth_rate/enrichment value; numerically ln(selected_fraction/input_fraction) for this table.
log2_enrichment
Derived log2(selected_fraction/input_fraction), calculated from the raw counts with a pseudocount of one.

Quality control

The authors counted occurrences of known native fragments in input and selected populations, added a pseudocount of one, normalized counts by library size, and calculated enrichment. The GEO table is already post-processing: all 11856 retained rows have unique A/C/G/T sequences of length 2–50 nt, nonnegative integer counts, and finite enrichment values, so no additional rows were removed. The 1327 fragments with zero input counts are retained because the published pseudocount procedure gives them finite scores; they should be treated as lower-coverage measurements. The deposited table contains 11856 detected fragments from a reported 11962-fragment library.

Curation notes

This is a pooled plasmid reporter selection and is included as an MPRA-like assay even though it does not use a conventional independent barcode RNA/DNA ratio. The source UTR_name field is retained because it links each fragment to its native yeast gene/fragment origin. The paper’s model-validation notebook uses a t0 >100 subset for a high-confidence prediction test, but that is not the authors’ library-wide assay QC and was not imposed on this complete published measurement table. The aggregate GEO file has no replicate-level counts or raw FASTQ files.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.