Modern deep learning in bioinformatics.

Modern deep learning in bioinformatics.
复制标题

DOI:
10.1093/jmcb/mjaa030
复制
发表时间:
2020-10-30
影响因子:
5.5
通讯作者:
Gao X
Gao X
中科院分区:
生物学1区
文献类型:
--
作者:
Li H;Tian S;Li Y;Fang Q;Tan R;Pan Y;Huang C;Xu Y;Gao X

文献摘要

参考文献

被引文献

相似文献

ML一直是最近人工智能复苏的主要贡献者。现代机器学习技术中最重要的部分是深度学习。深度学习建立在人工神经网络(ann)的基础上,理论上已被证明能够在任何指定的精度范围内近似任何非线性函数(Hornik, 1991),并已广泛用于解决各种计算任务(Li et al., 2019)。然而,它们被批评为黑盒子。这种可解释性的缺乏限制了它们的应用,特别是当它们的性能在其他更具可解释性的ML方法(如线性回归、逻辑回归、支持向量机和决策树)中无法脱颖而出时。在过去的十年中,科学技术的三个重要进步导致了人工神经网络的复兴,特别是通过深度学习。首先,现代生活中产生了空前数量的数据,主要是图像和自然语言数据。从这些数据中提取信息的复杂性给其他机器学习方法带来了巨大的挑战,但人工神经网络已经很好地处理了这些问题。同样,高通量生物学数据,如下一代测序、代谢组学数据、蛋白质组学数据和电子显微镜结构数据,也提出了同样具有挑战性的计算问题。其次,计算能力一直在以可承受的成本快速增长,包括新的计算设备的发展,如图形处理单元和现场可编程门阵列。这些设备为高度并行的模型提供了理想的硬件平台。第三,与大数据时代的竞争技术相比,一系列提出的优化算法使深度人工神经网络成为大型复杂数据分析和信息发现的理想技术。生物信息学领域还存在一些亟待解决的问题:首先,模型的可解释性对于生物学家理解模型如何帮助解决生物学问题至关重要,例如预测dna -蛋白质结合(Luo et al., 2020)。其次,与医疗保健或疾病诊断相关的计算模型的临床期望准确率为98%-99%,很难达到这么高的准确率。此外,两个
ML has been the main contributor to the recent resurgence of artificial intelligence. The most essential piece in modern ML technology is DL. DL is founded on artificial neural networks (ANNs), which have been theoretically proven to be capable of approximating any nonlinear function within any specified accuracy (Hornik, 1991) and have been widely used to solve various computational tasks (Li et al., 2019). However, they have been criticized for being black boxes. This lack of interpretability has limited their applications, particularly when their performance did not stand out among other more interpretable ML methods, such as linear regression, logistic regression, support vector machines, and decision trees. During the past decade, three important advances in science and technology have led to the rejuvenation of ANNs, particularly via DL. First, unprecedented quantities of data have been generated in modern life, mostly imaging and natural language data. The complex nature of information derivation from such data has posed great challenges to other ML methods but has been handled well by ANNs. Similarly, high-throughput biological data such as next-generation sequencing, metabolomic data, proteome data, and electron microscopic structural data, has raised equally challenging computational problems. Second, computational power has been increasing rapidly with affordable costs, including the development of new computing devices, such as graphics processing units and field programmable gate arrays. Such devices provide ideal hardware platforms for highly parallel models. Third, a range of proposed optimization algorithms have made deep ANNs stand out as an ideal technique for large and complex data analyses and information discovery compared to competing techniques in the big data era. Here are also some problems in the bioinformatics field as follows, which need to be tackled. First, the interpretability of model is essential to biologists to understand how model helps solve the biological problem, eg predicting DNA–protein binding (Luo et al., 2020). Second, the clinical expect accuracy of computational model related to the healthcare or disease diagnosis is $98%–99% and it is tough to reach that high accuracy. Moreover, two
DEEPre:通过深度学习进行基于序列的酶 EC 数预测
DOI: 10.1093/bioinformatics/btx680
发表时间: 2018-03-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Li Y;Wang S;Umarov R;Xie B;Fan M;Li L;Gao X
通讯作者: Gao X
DOI: 10.1093/bioinformatics/btx275
发表时间: 2017-09-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Alshahrani M;Khan MA;Maddouri O;Kinjo AR;Queralt-Rosinach N;Hoehndorf R
通讯作者: Hoehndorf R
DOI: 10.1016/j.ymeth.2019.04.008
发表时间: 2019-08-15
期刊: METHODS
影响因子: 4.8
作者:
Li, Yu;Huang, Chao;Gao, Xin
通讯作者: Gao, Xin
DOI: 10.1016/0893-6080(91)90009-t
发表时间: 1991-01-01
期刊: NEURAL NETWORKS
影响因子: 7.8
作者:
HORNIK, K
通讯作者: HORNIK, K
DOI: 10.1038/nature14236
发表时间: 2015-02-26
期刊: NATURE
影响因子: 64.8
作者:
Mnih, Volodymyr;Kavukcuoglu, Koray;Hassabis, Demis
通讯作者: Hassabis, Demis