Targeted RNA-seq measured reporter RNA relative to plasmid DNA for the same episomal dual-fluorescence library of natural S. cerevisiae and S. paradoxus 5′ transcript leaders. Three biological replicates are reported as RNA/DNA activity values, with the matched FACS-seq library QC applied to keep the element set consistent.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Organism
Budding yeast
Taxonomy ID
NCBITaxon:4932
Biosample
UNMAPPED:Saccharomyces_cerevisiae_BY4741
Reference genome
Not reported / not applicable
Design focus
Region-focused
Region of interest
Not reported / not applicable
Perturbation & assay details
Basal / Untreated; unstressed log-phase growth at 30°C
The same ENO2-YFP-mCherry episomal reporter library was assayed by targeted RNA-seq. Reporter RNA and plasmid DNA reads were normalized as RNA RPKM/DNA RPKM; the workbook reports Rep1RNA, Rep2RNA, Rep3RNA, and MeanRNA. The designed ENO2 transcription start sites were used on average 97% of the time. The library contains natural 5′ TLs up to 180 nt from S. cerevisiae and S. paradoxus, selected for at least 10% representation of host-gene transcripts.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (39 of 39)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
Page 1 · 50 rows · More results available
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 39 definitions
element_id
Original @UTR/Construct identifier formatted as gene|source context;start;end|construct state.
gene_or_locus
Gene or locus label from the original element identifier.
source_context
Chromosome label for S. cerevisiae (chr...) or source-genome label for S. paradoxus (Spar_...).
source_species
Species from which the tested transcript-leader sequence was derived, inferred from source_context.
coordinate_start
Start coordinate encoded in the source element identifier; strand-specific order is preserved.
coordinate_end
End coordinate encoded in the source element identifier; strand-specific order is preserved.
construct_state
Construct state reported by the source; all retained rows are WT.
length_nt
Designed transcript-leader length in nucleotides.
uaug_count
Number of upstream AUGs in the transcript leader, as reported by the source workbook.
utr_sequence
Designed natural 5′ transcript-leader sequence in the reporter, written 5′ to 3′.
kozak_context
Reporter start-codon Kozak context supplied in the source workbook.
feature_data_available
Whether the element also occurs in the source workbook Feature Data table.
freq_a
Fraction of transcript-leader nucleotides that are A.
freq_t
Fraction of transcript-leader nucleotides that are T.
freq_g
Fraction of transcript-leader nucleotides that are G.
freq_c
Fraction of transcript-leader nucleotides that are C.
max_a_stretch
Longest consecutive A run in the transcript leader.
num_gggg_quartets
Number of GGGG quartets using the source definition (floor of each consecutive G run divided by four).
kozak_strength
Source Kozak-strength score for the reporter start context.
ddg_median
Median predicted ΔΔG for unfolding the structure around the main start codon.
ddg_avg
Mean predicted ΔΔG for unfolding the structure around the main start codon.
g_quadruplex_count
Predicted G-quadruplex count reported by the source feature table.
cap40nt_folding_dg
Predicted folding free energy for the first 40 nt near the 5′ cap; source field Cap40ntFodling.
cap40nt_folding_abs
Absolute value of the first-40-nt cap folding energy used for modeling.
lsm_kozak_start
Adjusted Kozak score from the leaky-scanning model.
u_max
Longest homopolymeric U/T run in the transcript leader; T is used because the source sequence is DNA-encoded.
cap_proximal_a
A fraction among the first ≤20 nt near the 5′ cap.
cap_proximal_c
C fraction among the first ≤20 nt near the 5′ cap.
cap_proximal_g
G fraction among the first ≤20 nt near the 5′ cap.
cap_proximal_t
T fraction among the first ≤20 nt near the 5′ cap.
distal_a
A fraction among the last ≤30 nt near the main start codon.
distal_c
C fraction among the last ≤30 nt near the main start codon.
distal_g
G fraction among the last ≤30 nt near the main start codon.
distal_t
T fraction among the last ≤30 nt near the main start codon.
replicate_1_rna_dna_ratio
Reporter RNA abundance divided by plasmid DNA abundance for RNA-seq replicate 1.
replicate_2_rna_dna_ratio
Reporter RNA abundance divided by plasmid DNA abundance for RNA-seq replicate 2.
replicate_3_rna_dna_ratio
Reporter RNA abundance divided by plasmid DNA abundance for RNA-seq replicate 3.
mean_rna_dna_ratio
Source mean of the three RNA/DNA replicate ratios.
rna_replicate_sd
Sample standard deviation of the three RNA/DNA replicate ratios.
Quality control
The paper does not specify a separate RNA-specific exclusion threshold. To provide a matched, reproducible reporter set, this package applies the paper's shared three-replicate FACS QC (YFP/mCherry standard deviation ≤0.05 and at least 50 normalized reads in every replicate) and requires numeric RNA/DNA values: 10,257 of 11,027 source rows retained. Zero RNA/DNA values are retained as valid low-abundance measurements rather than treated as missing data.
Curation notes
The assay host is S. cerevisiae BY4741 (NCBI BioSample SAMN18740588; SRA project PRJNA721222), while the library sequences include both S. cerevisiae and S. paradoxus transcript leaders. The source element IDs preserve strand-specific coordinate order and the paper does not state a reference assembly, so reference_genome is null. Feature columns are merged from source Table_3 when available (10,181 of 10,257 retained rows); assay values and sequences come from source Table_1. RNA/DNA values are observational reporter-abundance measurements, not allele contrasts or differential-expression p-values.