Study / S47DUX1V12024-02-17

Evaluation and optimization of sequence-based gene regulatory deep learning models

Abdul Muntakim Rafi, Daria Nogina, Dmitry Penzar, Dohoon Lee, Danyeong Lee et al.

About this study

Neural networks have emerged as immensely powerful tools in predicting functional genomic regions, notably evidenced by recent successes in deciphering gene regulatory logic. However, a systematic evaluation of how model architectures and training strategies impact genomics model performance is lacking. To address this gap, we held a DREAM Challenge where competitors trained models on a dataset of millions of random promoter DNA sequences and corresponding expression levels, experimentally determined in yeast, to best capture the relationship between regulatory DNA and gene expression. For a robust evaluation of the models, we designed a comprehensive suite of benchmarks encompassing various sequence types. While some benchmarks produced similar results across the top-performing models, others differed substantially. All top-performing models used neural networks, but diverged in architectures and novel training strategies, tailored to genomics sequence data. To dissect how architectural and training choices impact performance, we developed the Prix Fixe framework to divide any given model into logically equivalent building blocks. We tested all possible combinations for the top three models and observed performance improvements for each. The DREAM Challenge models not only achieved state-of-the-art results on our comprehensive yeast dataset but also consistently surpassed existing benchmarks on Drosophila and human genomic datasets. Overall, we demonstrate that high-quality gold-standard genomics datasets can drive significant progress in model development.

Full author list & citation

Abdul Muntakim Rafi, Daria Nogina, Dmitry Penzar, Dohoon Lee, Danyeong Lee, Nayeon Kim, Sangyeup Kim, Dohyeon Kim, Yeojin Shin, Il-Youp Kwak, Georgy Meshcheryakov, Andrey Lando, Arsenii Zinkevich, Byeong-Chan Kim, Juhyun Lee, Taein Kang, Eeshit Dhaval Vaishnav, Payman Yadollahpour, Random Promoter DREAM Challenge Consortium, Sun Kim, Jake Albrecht, Aviv Regev, Wuming Gong, Ivan V. Kulakovskiy, Pablo Meyer, Carl de Boer. Evaluation and optimization of sequence-based gene regulatory deep learning models. 2024-02-17. https://doi.org/10.1101/2023.04.26.538471

Experiments 1

E5F7ZAIU9

Random Promoter DREAM Challenge held-out test MPRA library

A separate, higher-depth yeast reporter experiment measured a 71,103-sequence test library spanning random, native yeast, high/low-expression, challenging, SNV, motif-perturbation, and motif-tiling constructs. The library contains 80-bp variable inserts in a fixed promoter-like YFP reporter context; the processed table retains the source MAUDE expression/activity score and source pair annotations.

Sort-Seq / Flow-Seq MPRANCBITaxon:559292
Explore data

Raw source data 4 files

Original supplemental and deposited inputs retained for this study. Download files individually or together as a ZIP; nested folders are preserved. Source reuse terms apply, and sequencing reads may be omitted.

Download all 4 files (ZIP)GSE254493_filtered_test_data_with_MAUDE_expression.txt.gzREADME.txttest_subset_ids.tar.gzzenodo_record_10633252.json

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.