Study / S0YRGQ9NZ2020-04-30
Model-driven generation of artificial yeast promoters
Benjamin J. Kotopka, Christina D. Smolke
About this study
Promoters play a central role in controlling gene regulation; however, a small set of promoters is used for most genetic construct design in the yeast Saccharomyces cerevisiae. Generating and utilizing models that accurately predict protein expression from promoter sequences would enable rapid generation of useful promoters and facilitate synthetic biology efforts in this model organism. We measure the gene expression activity of over 675,000 sequences in a constitutive promoter library and over 327,000 sequences in an inducible promoter library. Training an ensemble of convolutional neural networks jointly on the two data sets enables very high (R² > 0.79) predictive accuracies on multiple sequence-activity prediction tasks. We describe model-guided design strategies that yield large, sequence-diverse sets of promoters exhibiting activities higher than those represented in training data and similar to current best-in-class sequences. Our results show the value of model-guided design as an approach for generating useful DNA parts.
Full author list & citation
Benjamin J. Kotopka, Christina D. Smolke. Model-driven generation of artificial yeast promoters. 2020-04-30. https://doi.org/10.1038/s41467-020-15977-4
Experiments 2
E19UM2OMO
A two-color episomal reporter library tested synthetic ZEV motif-containing promoters in Saccharomyces cerevisiae strain CSY1252, which expresses the ZEV artificial transcription factor. The library was sorted into 12 fluorescence bins under 0 µM beta-estradiol and 1 µM beta-estradiol, producing paired uninduced and induced promoter activity estimates for each sequence.
E6NSIPPXL
A two-color episomal reporter library tested synthetic, motif-preserving PGPD promoters in Saccharomyces cerevisiae strain CSY3 (W303 MATα). The final 313-bp consensus sequences were sorted into 12 GFP:mCherry fluorescence bins in two independent replicate experiments and quantified by full-length MiSeq sequence assignment linked to high-depth NextSeq activity estimates.