Predicting the sequence specificities of DNA- and RNA-binding proteins by deep learning

Predicting the sequence specificities of DNA- and RNA-binding proteins by deep learning
复制标题

DOI:
10.1038/nbt.3300
复制
发表时间:
2015-08-01
影响因子:
46.9
通讯作者:
Frey, Brendan J.
Frey, Brendan J.
中科院分区:
工程技术1区
文献类型:
--
作者:
Alipanahi, Babak;Delong, Andrew;Frey, Brendan J.

文献摘要

被引文献

相似文献

了解DNA和RNA结合蛋白的序列特异性对于开发生物系统中的调节过程模型和识别致病性疾病变体至关重要。在这里,我们表明,序列特异性可以确定从实验数据与“深度学习”技术,它提供了一个可扩展的,灵活的和统一的计算方法的模式发现。使用各种实验数据和评估指标,我们发现深度学习的性能优于其他最先进的方法,即使在体外数据训练和体内数据测试时也是如此。我们将这种方法称为DeepBind,并构建了一个独立的软件工具,该工具是全自动的,每次实验可以处理数百万个序列。由DeepBind确定的特异性很容易被可视化为位置权重矩阵的加权集合或“突变图”,其指示变异如何影响特定序列内的结合。
Knowing the sequence specificities of DNA- and RNA-binding proteins is essential for developing models of the regulatory processes in biological systems and for identifying causal disease variants. Here we show that sequence specificities can be ascertained from experimental data with 'deep learning' techniques, which offer a scalable, flexible and unified computational approach for pattern discovery. Using a diverse array of experimental data and evaluation metrics, we find that deep learning outperforms other state-of-the-art methods, even when training on in vitro data and testing on in vivo data. We call this approach DeepBind and have built a stand-alone software tool that is fully automatic and handles millions of sequences per experiment. Specificities determined by DeepBind are readily visualized as a weighted ensemble of position weight matrices or as a 'mutation map' that indicates how variations affect binding within a specific sequence.