Predicting mRNA Abundance Directly from Genomic Sequence Using Deep Convolutional Neural Networks

Predicting mRNA Abundance Directly from Genomic Sequence Using Deep Convolutional Neural Networks
复制标题

DOI:
10.1016/j.celrep.2020.107663
复制
发表时间:
2020-05-19
期刊:
影响因子:
8.8
通讯作者:
Shendure, Jay
Shendure, Jay
中科院分区:
生物学1区
文献类型:
--
作者:
Agarwal, Vikram;Shendure, Jay

文献摘要

被引文献

相似文献

仅从一级序列准确预测基因结构的算法对于注释人类基因组是变革性的。我们是否也可以仅仅根据基因组序列来预测基因的表达水平?在这里,我们试图将深度卷积神经网络应用于这一目标。令人惊讶的是,仅包括启动子序列和与mRNA稳定性相关的特征的模型分别解释了人类和小鼠稳态mRNA水平的59%和71%的变化。这种模型称为XIP-seq,其准确性是基于序列的替代模型的两倍多,并且与依赖于染色免疫沉淀测序(ChIP-seq)数据的模型一样具有预测性。XcRNA概括了全基因组范围内的转录活性模式,其残差可用于量化增强子、异染色质结构域和microRNA的影响。模型解释表明,启动子近端CpG二核苷酸强烈预测转录活性。展望未来,我们提出仅基于一级序列的细胞类型特异性基因表达预测是该领域的一个重大挑战。
Algorithms that accurately predict gene structure from primary sequence alone were transformative for annotating the human genome. Can we also predict the expression levels of genes based solely on genome sequence? Here, we sought to apply deep convolutional neural networks toward that goal. Surprisingly, a model that includes only promoter sequences and features associated with mRNA stability explains 59% and 71% of variation in steady-state mRNA levels in human and mouse, respectively. This model, termed Xpresso, more than doubles the accuracy of alternative sequence-based models and isolates rules as predictive as models relying on chromatic immunoprecipitation sequencing (ChIP-seq) data. Xpresso recapitulates genome-wide patterns of transcriptional activity, and its residuals can be used to quantify the influence of enhancers, heterochromatic domains, and microRNAs. Model interpretation reveals that promoter-proximal CpG dinucleotides strongly predict transcriptional activity. Looking forward, we propose cell-type-specific gene-expression predictions based solely on primary sequences as a grand challenge for the field.