Study / S1B3K2UEC2019-06-06
A Deep Neural Network for Predicting and Engineering Alternative Polyadenylation
Nicholas Bogard, Johannes Linder, Alexander B. Rosenberg, Georg Seelig
About this study
Alternative polyadenylation (APA) is a major driver of transcriptome diversity in human cells. Here, we use deep learning to predict APA from DNA sequence alone. We trained our model (APARENT, APA REgression NeT) on isoform expression data from over three million APA reporters. APARENT’s predictions are highly accurate when tasked with inferring APA in synthetic and human 3’UTRs. Visualizing features learned across all network layers reveals that APARENT recognizes sequence motifs known to recruit APA regulators, discovers previously unknown sequence determinants of 3’-end processing, and integrates these features into a comprehensive, interpretable cis-regulatory code. We apply APARENT to forward engineer functional polyadenylation signals with precisely defined cleavage position and isoform usage and validate predictions experimentally. Finally, we use APARENT to quantify the impact of genetic variants on APA. Our approach detects pathogenic variants in a wide range of disease contexts, expanding our understanding of the genetic origins of disease.
Full author list & citation
Nicholas Bogard, Johannes Linder, Alexander B. Rosenberg, Georg Seelig. A Deep Neural Network for Predicting and Engineering Alternative Polyadenylation. 2019-06-06. https://doi.org/10.1016/j.cell.2019.04.046
Experiments 2
E53YJBUXU
Transient episomal minigene reporter libraries tested randomized or degenerate sequence contexts from human 3′ UTRs in HEK293T cells. RNA-seq-derived UMI counts quantify proximal, distal/non-proximal, and de novo cleavage outcomes for millions of reporters.
E9ZELDUS5
An episomal 3′ UTR APA reporter array measured ClinVar, ACMG/HGMD-associated, and saturation-mutagenesis SNVs in human polyadenylation-site contexts. The processed table preserves the repository’s measured variant effect, APARENT prediction, significance, and sequence annotations in one variant-level table.