A machine learning-based framework for modeling transcription elongation

A machine learning-based framework for modeling transcription elongation
复制标题

基于机器学习的转录延伸建模框架

DOI:
10.1073/pnas.2007450118
复制
发表时间:
2021-02-09
影响因子:
11.1
通讯作者:
Zeng, Jianyang
Zeng, Jianyang
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Feng, Peiyuan;Xiao, An;Zeng, Jianyang

文献摘要

被引文献

相似文献

RNA聚合酶II (RNA polymerase II, Pol II)通常在沿基因体的某些位置暂停,从而中断转录延伸过程,这一过程往往伴随着前体mRNA剪接和基因表达调控等多种重要的生物学功能。表征转录延伸动力学可以帮助我们理解真核细胞中许多基本的生物过程。然而,实验测量Pol II延伸率通常是耗时和资源消耗。我们通过基于注意力的深度神经网络开发了PEPMAN (polymerase II伸长暂停模型),这是一个基于深度学习的模型,可以根据原生伸长转录物测序(NET-seq)数据准确预测Pol II暂停位点。通过充分利用注意机制,PEPMAN能够破译Pol II暂停背后的重要序列特征。更重要的是,我们证明了对各种类型的备选剪接位点周围的pepman预测结果的分析可以为理解共转录剪接事件提供有用的线索。此外,将PEPMAN预测结果与不同的表观遗传特征相关联,有助于揭示与转录延伸过程相关的重要因素。这些结果表明,PEPMAN可以从现有的高通量测序数据中为转录延伸建模和了解相关生物学因素提供有用和有效的工具。
RNA polymerase II (Pol II) generally pauses at certain positions along gene bodies, thereby interrupting the transcription elongation process, which is often coupled with various important biological functions, such as precursor mRNA splicing and gene expression regulation. Characterizing the transcriptional elongation dynamics can thus help us understand many essential biological processes in eukaryotic cells. However, experimentally measuring Pol II elongation rates is generally time and resource consuming. We developed PEPMAN (polymerase II elongation pausing modeling through attention-based deep neural network), a deep learning-based model that accurately predicts Pol II pausing sites based on the native elongating transcript sequencing (NET-seq) data. Through fully taking advantage of the attention mechanism, PEPMAN is able to decipher important sequence features underlying Pol II pausing. More importantly, we demonstrated that the analyses of the PEPMAN-predicted results around various types of alternative splicing sites can provide useful clues into understanding the cotranscriptional splicing events. In addition, associating the PEPMAN prediction results with different epigenetic features can help reveal important factors related to the transcription elongation process. All these results demonstrated that PEPMAN can provide a useful and effective tool for modeling transcription elongation and understanding the related biological factors from available high-throughput sequencing data.