A separately measured library of designed 80-bp promoter constructs was assayed in S288C Δura3 yeast using the same episomal YFP/RFP GPRA workflow. The library includes model-designed sequences, natural and random controls, mutational trajectories, and motif or sequence perturbation constructs; the processed table uses the author's ≥100-read release.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Organism
Budding yeast
Taxonomy ID
NCBITaxon:4932
Biosample
UNMAPPED:S288CdeltaUra3
Reference genome
R64
Design focus
Synthetic / Motif-focused
Region of interest
Not reported / not applicable
Perturbation & assay details
Basal / SD-Ura defined medium
Designed single-stranded oligonucleotides from Twist Bioscience were cloned into the same low-copy CEN YFP/RFP reporter plasmid used for GPRA. S288C Δura3 yeast were grown in SD-Ura at approximately 30°C, sorted by the RFP:YFP ratio into 18 uniform bins, and promoter amplicons were sequenced as 2 × 76-bp reads. Designed reads were aligned to the ordered promoter sequences with Bowtie2, only perfect matches were used, and the author score is the read-count-weighted mean of the bin labels. variable_sequence is the nominal 80-bp sequence between the fixed reporter flanks.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (12 of 12)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
Page 1 · 50 rows · More results available
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 12 definitions
element_id
Package-generated identifier for a retained designed promoter construct.
source_id
Original origID label from the GEO designed-library release; may contain comma-separated aliases.
sequence_id
Original seqID label from the GEO designed-library release; may contain comma-separated aliases or semicolon-separated paired labels.
sequence
Full 110-nt designed promoter sequence, including fixed distal and proximal reporter-scaffold flanks.
variable_sequence
The 80-bp designed sequence between the fixed reporter flanks.
sequence_length
Length of sequence in nucleotides; all retained rows are 110.
variable_length
Length of variable_sequence in nucleotides; all retained rows are 80.
mean_expression
Author-provided expression level, calculated as a read-count-weighted mean across 18 FACS bins.
gc_fraction
Fraction of bases in sequence that are G or C.
variable_gc_fraction
Fraction of bases in variable_sequence that are G or C.
source_id_prefix
First dot-delimited token of source_id, retained as a compact design-family label without imposing a biological interpretation.
qc_status
Package QC status; all rows in the processed table passed the author-release and package filters.
Quality control
The authors aligned designed-library reads to the ordered sequences and retained only perfect matches. They calculated scores for all promoters with observed reads and used the ≥100-read release for publication analyses. Package QC retained rows with four tab-separated fields, nonempty source identifiers, an exactly 110-nt A/C/G/T-only sequence, and a finite numeric expression score; 67,928 of 68,086 rows in the min100Reads source passed. The 158 excluded rows had an NA expression score. The unfiltered all and min100Reads releases are both retained in raw_data.
Curation notes
The processed table is based on GSE163045_MolEvol_seq_data_SCUra_only.splitByOrigID.meanEL.min100Reads.txt.gz, retaining 67,928 scored constructs and excluding 158 NA-score rows. The separate all.txt.gz release contains 75,487 records and is retained for provenance. Source identifiers are quoted as valid CSV because several designed records contain comma-separated aliases. The library is heterogeneous and has no single genomic locus; source_id and source_id_prefix preserve the authors' design labels for grouping model-designed, trajectory, random-control, native-promoter, and perturbation constructs. Pairwise variant effects are not re-estimated because the GEO release supplies sequence-level weighted-bin scores rather than per-construct count matrices.