Experiment / E64CXX7C6Sort-Seq / Flow-Seq MPRA

FACS-seq protein output from natural yeast transcript leaders

Deciphering the landscape of cis-acting sequences in natural yeast transcript leaders

Three biological replicates of an episomal dual-fluorescence reporter library tested natural 5′ transcript leaders from Saccharomyces cerevisiae and Saccharomyces paradoxus upstream of YFP, with mCherry as an internal control. Cells were FACS-binned by the YFP/mCherry ratio and construct counts were used to estimate normalized protein output for each transcript leader.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated; unstressed log-phase growth at 30°C

The ENO2 promoter drove the reporter transcript, and each natural 5′ TL was cloned upstream of YFP in a dual YFP/mCherry plasmid. Cells were sorted into eight expression bins by YFP/mCherry; bin counts were normalized to the proportion of cells sorted, and mean YFP/mCherry was calculated. The library contains 11,027 wild-type TLs up to 180 nt from S. cerevisiae and S. paradoxus, selected for at least 10% representation of host-gene transcripts. The assay directly sequenced reporter constructs rather than using a separate barcode field.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (43 of 43)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 43 definitions
element_id
Original @UTR/Construct identifier formatted as gene|source context;start;end|construct state.
gene_or_locus
Gene or locus label from the original element identifier.
source_context
Chromosome label for S. cerevisiae (chr...) or source-genome label for S. paradoxus (Spar_...).
source_species
Species from which the tested transcript-leader sequence was derived, inferred from source_context.
coordinate_start
Start coordinate encoded in the source element identifier; strand-specific order is preserved.
coordinate_end
End coordinate encoded in the source element identifier; strand-specific order is preserved.
construct_state
Construct state reported by the source; all retained rows are WT.
length_nt
Designed transcript-leader length in nucleotides.
uaug_count
Number of upstream AUGs in the transcript leader, as reported by the source workbook.
utr_sequence
Designed natural 5′ transcript-leader sequence in the reporter, written 5′ to 3′.
kozak_context
Reporter start-codon Kozak context supplied in the source workbook.
feature_data_available
Whether the element also occurs in the source workbook Feature Data table.
freq_a
Fraction of transcript-leader nucleotides that are A.
freq_t
Fraction of transcript-leader nucleotides that are T.
freq_g
Fraction of transcript-leader nucleotides that are G.
freq_c
Fraction of transcript-leader nucleotides that are C.
max_a_stretch
Longest consecutive A run in the transcript leader.
num_gggg_quartets
Number of GGGG quartets using the source definition (floor of each consecutive G run divided by four).
kozak_strength
Source Kozak-strength score for the reporter start context.
ddg_median
Median predicted ΔΔG for unfolding the structure around the main start codon.
ddg_avg
Mean predicted ΔΔG for unfolding the structure around the main start codon.
g_quadruplex_count
Predicted G-quadruplex count reported by the source feature table.
cap40nt_folding_dg
Predicted folding free energy for the first 40 nt near the 5′ cap; source field Cap40ntFodling.
cap40nt_folding_abs
Absolute value of the first-40-nt cap folding energy used for modeling.
lsm_kozak_start
Adjusted Kozak score from the leaky-scanning model.
u_max
Longest homopolymeric U/T run in the transcript leader; T is used because the source sequence is DNA-encoded.
cap_proximal_a
A fraction among the first ≤20 nt near the 5′ cap.
cap_proximal_c
C fraction among the first ≤20 nt near the 5′ cap.
cap_proximal_g
G fraction among the first ≤20 nt near the 5′ cap.
cap_proximal_t
T fraction among the first ≤20 nt near the 5′ cap.
distal_a
A fraction among the last ≤30 nt near the main start codon.
distal_c
C fraction among the last ≤30 nt near the main start codon.
distal_g
G fraction among the last ≤30 nt near the main start codon.
distal_t
T fraction among the last ≤30 nt near the main start codon.
replicate_1_yfp_mcherry
Normalized YFP/mCherry protein-output estimate for FACS-seq replicate 1.
replicate_2_yfp_mcherry
Normalized YFP/mCherry protein-output estimate for FACS-seq replicate 2.
replicate_3_yfp_mcherry
Normalized YFP/mCherry protein-output estimate for FACS-seq replicate 3.
mean_yfp_mcherry
Source mean of the three normalized YFP/mCherry replicate estimates.
yfp_replicate_sd
Sample standard deviation of the three YFP/mCherry replicate estimates; used for QC.
replicate_1_normalized_reads
Normalized FACS-bin sequencing read count for replicate 1.
replicate_2_normalized_reads
Normalized FACS-bin sequencing read count for replicate 2.
replicate_3_normalized_reads
Normalized FACS-bin sequencing read count for replicate 3.
min_replicate_normalized_reads
Minimum normalized FACS-bin read count across the three replicates; used for QC.

Quality control

The paper removed noisy TLs with three-replicate YFP/mCherry standard deviation >0.05 or fewer than 50 normalized reads. This package retains rows with replicate standard deviation ≤0.05 and at least 50 normalized reads in every replicate: 10,257 of 11,027 source rows. No source row exceeded the standard-deviation threshold; 770 rows were excluded by the minimum read-depth rule.

Curation notes

The assay host is S. cerevisiae BY4741 (NCBI BioSample SAMN18740588; SRA project PRJNA721222), while the library sequences include both S. cerevisiae and S. paradoxus transcript leaders. The source element IDs preserve strand-specific coordinate order and the paper does not state a reference assembly, so reference_genome is null. Feature columns are merged from source Table_3 when available (10,181 of 10,257 retained rows). The source assay reports normalized protein output rather than a log2 allele effect, and all retained constructs are wild-type natural TLs.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.