N50-C random 3′ UTR growth-selection MPRA
Effects of sequence motifs in the yeast 3′ untranslated region determined from massively parallel assays of random sequencesA low-copy centromeric episomal reporter library of random 50-bp synthetic sequences was inserted into the 3′ UTR of a HIS3 reporter in the CYC1 promoter/terminator framework. The N50-C design replaced the first 151 bp of the CYC1 3′ UTR while retaining the cleavage site and downstream constant sequence; library-member abundance before versus after 3-AT growth selection provided the expression proxy.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
1 mM 3-amino-1,2,4-triazole (3-AT) in SD-Leu-His growth selection; SD-Leu culture was the pre-selection input
The p415-CYC1-HIS3 low-copy centromeric plasmid carried a LEU2 marker and a random 50-bp N50 sequence in the reporter 3′ UTR. N50-C removed the first 151 bp of the native CYC1 terminator and retained its cleavage site plus downstream constant sequence. BY4741 his3::KanMX yeast were grown before selection and to OD660 approximately 1 in SD-Leu-His plus 1 mM 3-AT; pre- and post-selection plasmid amplicons were sequenced on an Illumina NextSeq 550, clustered with Bartender, and aligned with Bowtie2. The read-frequency change is a continuous proxy for HIS3 protein expression and fitness, not a direct RNA/DNA stability measurement.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (22 of 22)
| Row | ||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ||||||||||||||||||||||
| 2 | ||||||||||||||||||||||
| 3 | ||||||||||||||||||||||
| 4 | ||||||||||||||||||||||
| 5 | ||||||||||||||||||||||
| 6 | ||||||||||||||||||||||
| 7 | ||||||||||||||||||||||
| 8 | ||||||||||||||||||||||
| 9 | ||||||||||||||||||||||
| 10 | ||||||||||||||||||||||
| 11 | ||||||||||||||||||||||
| 12 | ||||||||||||||||||||||
| 13 | ||||||||||||||||||||||
| 14 | ||||||||||||||||||||||
| 15 | ||||||||||||||||||||||
| 16 | ||||||||||||||||||||||
| 17 | ||||||||||||||||||||||
| 18 | ||||||||||||||||||||||
| 19 | ||||||||||||||||||||||
| 20 | ||||||||||||||||||||||
| 21 | ||||||||||||||||||||||
| 22 | ||||||||||||||||||||||
| 23 | ||||||||||||||||||||||
| 24 | ||||||||||||||||||||||
| 25 | ||||||||||||||||||||||
| 26 | ||||||||||||||||||||||
| 27 | ||||||||||||||||||||||
| 28 | ||||||||||||||||||||||
| 29 | ||||||||||||||||||||||
| 30 | ||||||||||||||||||||||
| 31 | ||||||||||||||||||||||
| 32 | ||||||||||||||||||||||
| 33 | ||||||||||||||||||||||
| 34 | ||||||||||||||||||||||
| 35 | ||||||||||||||||||||||
| 36 | ||||||||||||||||||||||
| 37 | ||||||||||||||||||||||
| 38 | ||||||||||||||||||||||
| 39 | ||||||||||||||||||||||
| 40 | ||||||||||||||||||||||
| 41 | ||||||||||||||||||||||
| 42 | ||||||||||||||||||||||
| 43 | ||||||||||||||||||||||
| 44 | ||||||||||||||||||||||
| 45 | ||||||||||||||||||||||
| 46 | ||||||||||||||||||||||
| 47 | ||||||||||||||||||||||
| 48 | ||||||||||||||||||||||
| 49 | ||||||||||||||||||||||
| 50 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 22 definitions
- element_id
- Deterministic package identifier for a library member; the authors supplied the sequence rather than a separate element ID.
- source_row
- 1-based data-row number in the corresponding source CSV, excluding its header.
- library
- Source library context: N50-C or N50-EPC.
- sequence_dna
- 50-nt synthetic N50 sequence in the DNA alphabet; T is the DNA representation of RNA U.
- sequence_length_nt
- Length of the synthetic N50 sequence in nucleotides.
- input_count
- Authors' pre-selection sequencing read count for the library member.
- output_count
- Authors' post-selection sequencing read count for the library member.
- input_frequency
- Authors' pre-selection population frequency, retained exactly as supplied without rescaling.
- output_frequency
- Authors' post-selection population frequency, retained exactly as supplied without rescaling.
- log2_enrichment
- Authors' log2_score, defined as log2(post-selection frequency / pre-selection frequency); the growth-selection expression proxy.
- au_fraction
- Fraction of the 50-nt DNA sequence that is A or T, equivalent to A/U content in the RNA sequence.
- gc_fraction
- Fraction of the 50-nt DNA sequence that is G or C.
- contains_UAUAUA
- Boolean for the RNA efficiency-element motif UAUAUA, searched as DNA TATATA.
- contains_U5AUA
- Boolean for the RNA motif U[5]AUA (UUUUUAUA), searched as DNA TTTTTATA.
- contains_AAWAAA
- Boolean for the yeast positioning-element consensus AAWAAA, with W=A/U; searched as DNA AA[AT]AAA.
- contains_U8
- Boolean for an uninterrupted poly(U)8 motif, searched as DNA TTTTTTTT.
- contains_GCGCGC
- Boolean for the GC-rich control motif GCGCGC.
- contains_Puf1_Puf2_UAAUNNNUAAU
- Boolean for the Puf1/Puf2 motif UAAUNNNUAAU with N as any base, searched in DNA notation.
- contains_Puf3_UGUANAUA
- Boolean for the Puf3 motif UGUANAUA with N as any base, searched in DNA notation.
- contains_Puf4_UGUANANUA
- Boolean for the Puf4 motif UGUANANUA with N as any base, searched in DNA notation.
- contains_Puf5_UGUANNNNUA
- Boolean for the Puf5 motif UGUANNNNUA with N as any base, searched in DNA notation.
- contains_Puf6_UUGU
- Boolean for the Puf6 motif UUGU, searched as DNA TTGT.
Quality control
The authors required at least 5 reads in the input sample and at least 1 read in the output sample, with no pseudocounting; this produced 590,024 N50-C variants in the source table. Independent package QC verified 50-nt sequences, an A/C/G/T-only alphabet, unique sequence identities, positive finite counts/frequencies, finite log2 enrichment, and consistency of the source enrichment with log2(output_frequency/input_frequency) within 0.02. All 590,024 source rows passed and were retained.
Curation notes
This library was measured in one pooled growth-selection assay without biological replicates because a replicate transformation would generate a different random library. The sequence itself is the variant identity; there is no separate barcode or allele contrast. N50-C lacks invariant efficiency and positioning elements, so its score distribution is not directly interchangeable with N50-EPC without accounting for library context. Motif flags are derived annotations using DNA T→RNA U conversion and the N/W wildcard conventions stated in the paper. The processed table contains 590024 rows, including the authors’ low- and high-enrichment variants; the raw sequencing reads remain in BioProject PRJNA750726 and are not packaged.