A high-complexity library of nominally 80-bp random DNA inserts was placed between fixed promoter-scaffold sequences upstream of YFP in a low-copy episomal reporter and assayed in Saccharomyces cerevisiae strain Y8205. Cells were sorted into 18 FACS expression bins, and the author-released weighted-bin promoter scores are packaged with a deterministic sequence-level sample for the large source release.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Organism
Budding yeast
Taxonomy ID
NCBITaxon:4932
Biosample
UNMAPPED:Y8205
Reference genome
R64
Design focus
Synthetic / Motif-focused
Region of interest
Not reported / not applicable
Perturbation & assay details
Basal / SD-Ura defined medium
The GPRA construct was the low-copy CEN plasmid Addgene 127546, with a constitutive RFP reporter for extrinsic-noise control and YFP driven by a promoter containing the variable sequence. Yeast were grown in SD-Ura at approximately 30°C in log phase, sorted by the RFP:YFP ratio into 18 adjacent uniform bins, regrown for plasmid recovery, and sequenced as 2 × 76-bp promoter amplicons on an Illumina NextSeq 500. The source promoter score is the read-count-weighted mean of the 18 sorting-bin labels. The fixed sequence scaffold is 17 nt distal (TGCATTTTTTTCACATC) and 13 nt proximal (GGTTACGGCTGTT); variable_sequence is extracted between these flanks.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (10 of 10)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
Page 1 · 50 rows · More results available
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 10 definitions
element_id
Package-generated identifier for a sampled random promoter construct.
source_row_index
Zero-based line index in the headerless complete GEO random-library source file.
sequence
Observed full promoter sequence, including the fixed distal and proximal reporter-scaffold flanks.
variable_sequence
Observed sequence between the 17-nt distal and 13-nt proximal fixed flanks; nominally the 80-bp random insert.
sequence_length
Length of the observed full promoter sequence in nucleotides.
variable_length
Length of the observed variable sequence in nucleotides.
mean_expression
Author-provided expression level, calculated as a read-count-weighted mean across 18 FACS bins; higher values indicate greater reporter expression.
gc_fraction
Fraction of bases in sequence that are G or C.
variable_gc_fraction
Fraction of bases in variable_sequence that are G or C.
qc_status
Package QC status; all rows in the processed table passed the stated source-format and sequence filters.
Quality control
Author processing aligned paired reads with a 40 ± 15-bp central overlap, discarded reads failing that constraint, clustered related random promoters with Bowtie2, retained the highest-read representative for each cluster, and calculated expression as the weighted mean across 18 bins. Package QC retained records with exactly two tab-separated fields, a finite numeric score, an A/C/G/T-only sequence, and the expected fixed flanks; 829 ambiguous-base records were excluded. The author release contains observed collapsed sequence lengths of 97–127 nt, so no exact 110-nt filter was imposed. Because the complete source release is 668 MB compressed, table.csv is a deterministic every-25th-valid-source-row sample; the complete release is retained in raw_data.
Curation notes
GEO GSE163045 contains 21,037,407 rows in the current complete SD-Ura random-library release, while the paper's methods report 20,616,659 random sequences used for the defined-medium training model; the release was preserved in source order without attempting to reconcile that count difference. The complete compressed source is the authoritative sequence-expression archive. The processed table includes 841,452 valid rows (approximately 4% of the source) with full sequences, variable inserts, author scores, and source line indices so the sample can be traced back to raw_data. Raw FASTQ reads were not included.