Study / S8OFAY5NE2024-10-11

A community effort to optimize sequence-based deep learning models of gene regulation

Abdul Muntakim Rafi, Daria Nogina, Dmitry Penzar, Dohoon Lee, Danyeong Lee et al.

About this study

A systematic evaluation of how model architectures and training strategies impact genomics model performance is needed. To address this gap, we held a DREAM Challenge where competitors trained models on a dataset of millions of random promoter DNA sequences and corresponding expression levels, experimentally determined in yeast. For a robust evaluation of the models, we designed a comprehensive suite of benchmarks encompassing various sequence types. All top-performing models used neural networks but diverged in architectures and training strategies. To dissect how architectural and training choices affect performance, we developed the Prix Fixe framework to divide models into modular building blocks. We tested all possible combinations for the top three models, further improving their performance. The DREAM Challenge models not only achieved state-of-the-art results on our comprehensive yeast dataset but also consistently surpassed existing benchmarks on Drosophila and human genomic datasets, demonstrating the progress that can be driven by gold-standard genomics datasets.

Full author list & citation

Abdul Muntakim Rafi, Daria Nogina, Dmitry Penzar, Dohoon Lee, Danyeong Lee, Nayeon Kim, Sangyeup Kim, Dohyeon Kim, Yeojin Shin, Il-Youp Kwak, Georgy Meshcheryakov, Andrey Lando, Arsenii Zinkevich, Byeong-Chan Kim, Juhyun Lee, Taein Kang, Eeshit Dhaval Vaishnav, Payman Yadollahpour, Random Promoter DREAM Challenge Consortium, Sun Kim, Jake Albrecht, Aviv Regev, Wuming Gong, Ivan V. Kulakovskiy, Pablo Meyer, Carl G. de Boer, Susanne Bornelöv, Fredrik Svensson, Maria-Anna Trapotsi, Duc Tran, Tin Nguyen, Xinming Tu, Wuwei Zhang, Wei Qiu, Rohan Ghotra, Yiyang Yu, Ethan Labelson, Aayush Prakash, Ashwin Narayanan, Peter Koo, Xiaoting Chen, David T. Jones, Michele Tinti, Yuanfang Guan, Maolin Ding, Ken Chen, Yuedong Yang, Ke Ding, Gunjan Dixit, Jiayu Wen, Zhihan Zhou, Pratik Dutta, Rekha Sathian, Pallavi Surana, Yanrong Ji, Han Liu, Ramana V. Davuluri, Yu Hiratsuka, Mao Takatsu, Tsai-Min Chen, Chih-Han Huang, Hsuan-Kai Wang, Edward S. C. Shih, Sz-Hau Chen, Chih-Hsun Wu, Jhih-Yu Chen, Kuei-Lin Huang, Ibrahim Alsaggaf, Patrick Greaves, Carl Barton, Cen Wan, Nicholas Abad, Cindy Körner, Lars Feuerbach, Benedikt Brors, Yichao Li, Sebastian Röner, Pyaree Mohan Dash, Max Schubach, Onuralp Soylemez, Andreas Møller, Gabija Kavaliauskaite, Jesper Madsen, Zhixiu Lu, Owen Queen, Ashley Babjac, Scott Emrich, Konstantinos Kardamiliotis, Konstantinos Kyriakidis, Andigoni Malousi, Ashok Palaniappan, Krishnakant Gupta, Prasanna Kumar S, Jake Bradford, Dimitri Perrin, Robert Salomone, Carl Schmitz, Chen JiaXing, Wang JingZhe, Yang AiWei. A community effort to optimize sequence-based deep learning models of gene regulation. 2024-10-11. https://doi.org/10.1038/s41587-024-02414-w

Experiments 2

E8NQ46DKX

Random N80 yeast promoter library MPRA

A high-complexity library of nominally 80-bp random DNA inserts was placed between fixed distal and proximal promoter scaffold sequences upstream of YFP in a dual reporter and assayed in S288C ΔURA3 yeast. The package combines the author-released Zenodo train and validation sequence-expression tables for the random N80 library.

Sort-Seq / Flow-Seq MPRABudding yeastR64
Explore data
E9OHYUNS7

Designed NBT yeast promoter MPRA

A separately measured designed promoter library of approximately 71,000 constructs was assayed in S288C ΔURA3 yeast using the same dual-reporter FACS-bin workflow as the random library. The library contains native yeast fragments, random controls, high- and low-expression designs, challenging sequences, SNV perturbations, and motif perturbation or tiling constructs.

Sort-Seq / Flow-Seq MPRABudding yeastR64
Explore data

Raw source data 5 files

Original supplemental and deposited inputs retained for this study. Download files individually or together as a ZIP; nested folders are preserved. Source reuse terms apply, and sequencing reads may be omitted.

Download all 5 files (ZIP)GSE254493_filtered_test_data_with_MAUDE_expression.txt.gzGSE254493_series.txtzenodo_10633252_test_subset_ids.tar.gzzenodo_10633252_train.txtzenodo_10633252_val.txt

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.