Predicting human protein function with multi-task deep neural networks.

Predicting human protein function with multi-task deep neural networks.
复制标题

通过多任务深度神经网络预测人类蛋白质功能。

DOI:
10.1371/journal.pone.0198216
复制
发表时间:
2018
期刊:
影响因子:
3.7
通讯作者:
Jones DT
Jones DT
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Fa R;Cozzetto D;Wan C;Jones DT

文献摘要

参考文献

被引文献

相似文献

迫切需要机器学习方法来预测蛋白质功能,特别是现在,尽管基于序列相似性的功能分配被广泛使用,但很大一部分已知序列仍然没有被注释。监督学习在蛋白质功能预测中面临的一个主要瓶颈是问题的结构化、多标签性质,因为生物角色是由来自分层组织的受控词汇(如基因本体论)的术语列表表示的。在这项工作中,我们在深度学习领域的最新发展的基础上,研究了多任务深度神经网络(MTDNN)的有用性,该网络由上游共享层组成,这些共享层上并行堆叠了许多独立模块(具有自己的输出单元的附加隐层)作为输出GO项(任务)的数量。MTDNN学习单独的任务,部分使用共享表示,部分来自任务特定的特征。当不能识别与实验验证的功能相近的同源时,MTDNN给出了比基于公共数据库中注释频率或同源转移的基线方法更准确的预测。更重要的是,结果表明,MTDNN的二进制分类精度高于其他基于机器学习的方法,这些方法没有利用预测任务之间的共性和差异。有趣的是,与单任务预测器相比,MTDNN中的性能改进与任务数量并不是线性相关的,但在我们的情况下,中等规模的模型提供了更多的改进。MTDNN的优点之一是,给定一组特征,MTDNN不需要像传统机器学习算法那样具有引导特征选择过程。实验结果表明,提出的MTDNN算法提高了蛋白质功能预测的性能。另一方面,深度学习技术在进一步提升预测能力方面仍有很大的空间。
Machine learning methods for protein function prediction are urgently needed, especially now that a substantial fraction of known sequences remains unannotated despite the extensive use of functional assignments based on sequence similarity. One major bottleneck supervised learning faces in protein function prediction is the structured, multi-label nature of the problem, because biological roles are represented by lists of terms from hierarchically organised controlled vocabularies such as the Gene Ontology. In this work, we build on recent developments in the area of deep learning and investigate the usefulness of multi-task deep neural networks (MTDNN), which consist of upstream shared layers upon which are stacked in parallel as many independent modules (additional hidden layers with their own output units) as the number of output GO terms (the tasks). MTDNN learns individual tasks partially using shared representations and partially from task-specific characteristics. When no close homologues with experimentally validated functions can be identified, MTDNN gives more accurate predictions than baseline methods based on annotation frequencies in public databases or homology transfers. More importantly, the results show that MTDNN binary classification accuracy is higher than alternative machine learning-based methods that do not exploit commonalities and differences among prediction tasks. Interestingly, compared with a single-task predictor, the performance improvement is not linearly correlated with the number of tasks in MTDNN, but medium size models provide more improvement in our case. One of advantages of MTDNN is that given a set of features, there is no requirement for MTDNN to have a bootstrap feature selection procedure as what traditional machine learning algorithms do. Overall, the results indicate that the proposed MTDNN algorithm improves the performance of protein function prediction. On the other hand, there is still large room for deep learning techniques to further enhance prediction ability.
DOI: 10.1371/journal.pone.0063754
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Minneci F;Piovesan D;Cozzetto D;Jones DT
通讯作者: Jones DT
基因本体论:2011年的增强。
DOI: 10.1093/nar/gkr1028
发表时间: 2012-01
影响因子: 14.9
作者:
Gene Ontology Consortium
通讯作者: Gene Ontology Consortium
DOI: 10.1186/1471-2105-14-248
发表时间: 2013-08-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Hauser M;Mayer CE;Söding J
通讯作者: Söding J
DOI: 10.1186/s13059-016-1037-6
发表时间: 2016-09-07
期刊: Genome biology
影响因子: 12.3
作者:
Jiang Y;Oron TR;Clark WT;Bankapur AR;D'Andrea D;Lepore R;Funk CS;Kahanda I;Verspoor KM;Ben-Hur A;Koo da CE;Penfold-Brown D;Shasha D;Youngs N;Bonneau R;Lin A;Sahraeian SM;Martelli PL;Profiti G;Casadio R;Cao R;Zhong Z;Cheng J;Altenhoff A;Skunca N;Dessimoz C;Dogan T;Hakala K;Kaewphan S;Mehryary F;Salakoski T;Ginter F;Fang H;Smithers B;Oates M;Gough J;Törönen P;Koskinen P;Holm L;Chen CT;Hsu WL;Bryson K;Cozzetto D;Minneci F;Jones DT;Chapman S;Bkc D;Khan IK;Kihara D;Ofer D;Rappoport N;Stern A;Cibrian-Uhalte E;Denny P;Foulger RE;Hieta R;Legge D;Lovering RC;Magrane M;Melidoni AN;Mutowo-Meullenet P;Pichler K;Shypitsyna A;Li B;Zakeri P;ElShal S;Tranchevent LC;Das S;Dawson NL;Lee D;Lees JG;Sillitoe I;Bhat P;Nepusz T;Romero AE;Sasidharan R;Yang H;Paccanaro A;Gillis J;Sedeño-Cortés AE;Pavlidis P;Feng S;Cejuela JM;Goldberg T;Hamp T;Richter L;Salamov A;Gabaldon T;Marcet-Houben M;Supek F;Gong Q;Ning W;Zhou Y;Tian W;Falda M;Fontana P;Lavezzo E;Toppo S;Ferrari C;Giollo M;Piovesan D;Tosatto SC;Del Pozo A;Fernández JM;Maietta P;Valencia A;Tress ML;Benso A;Di Carlo S;Politano G;Savino A;Rehman HU;Re M;Mesiti M;Valentini G;Bargsten JW;van Dijk AD;Gemovic B;Glisic S;Perovic V;Veljkovic V;Veljkovic N;Almeida-E-Silva DC;Vencio RZ;Sharan M;Vogel J;Kansakar L;Zhang S;Vucetic S;Wang Z;Sternberg MJ;Wass MN;Huntley RP;Martin MJ;O'Donovan C;Robinson PN;Moreau Y;Tramontano A;Babbitt PC;Brenner SE;Linial M;Orengo CA;Rost B;Greene CS;Mooney SD;Friedberg I;Radivojac P
通讯作者: Radivojac P
DOI: 10.1186/1477-5956-7-27
发表时间: 2009-08-09
期刊: Proteome science
影响因子: 2
作者:
Lee BJ;Shin MS;Oh YJ;Oh HS;Ryu KH
通讯作者: Ryu KH