FFPred 2.0: improved homology-independent prediction of gene ontology terms for eukaryotic protein sequences.

FFPred 2.0: improved homology-independent prediction of gene ontology terms for eukaryotic protein sequences.
复制标题

DOI:
10.1371/journal.pone.0063754
复制
发表时间:
2013
期刊:
影响因子:
3.7
通讯作者:
Jones DT
Jones DT
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Minneci F;Piovesan D;Cozzetto D;Jones DT

文献摘要

参考文献

被引文献

相似文献

为了充分了解细胞行为,生物学家正在对人类基因组中的功能元素进行编目,并描述它们在各种组织和条件下的作用。然而,大约30%的人类蛋白质的功能信息——无论是实验验证的还是通过相似性计算推断的——仍然完全缺失。FFPred最初是为了填补这一空白而开发的,通过靶向具有已知功能的远同源或无同源的序列,并通过利用与特定分子活动和生物过程相关的内在紊乱的明确模式。在这里,我们提出了一个更新和改进的版本,它建立在更大的蛋白质序列和注释数据集上,并使用更新的成分特征预测器以及修订的训练程序。FFPred 2.0包含用于预测442个基因本体(GO)术语的支持向量回归模型,这在很大程度上扩展了本体的覆盖范围,特别是生物过程类别。氧化石墨烯术语表主要围绕大分子相互作用及其在调控、信号传导、发育和代谢过程中的作用展开。对新标注蛋白的基准实验表明,FFPred 2.0比其前身和ProtFun服务器提供了更准确的功能分配;此外,它的分配可以补充使用基于blast的注释转移获得的信息,特别是在生物过程类别的预测方面。此外,FFPred 2.0可以用于注释属于几种真核生物的蛋白质,但预测质量有一定的下降。我们通过使用精度召回图和COGIC分数来说明所有这些点,COGIC分数是我们最近提出的函数预测精度的替代数值评估指标。
To understand fully cell behaviour, biologists are making progress towards cataloguing the functional elements in the human genome and characterising their roles across a variety of tissues and conditions. Yet, functional information – either experimentally validated or computationally inferred by similarity – remains completely missing for approximately 30% of human proteins. FFPred was initially developed to bridge this gap by targeting sequences with distant or no homologues of known function and by exploiting clear patterns of intrinsic disorder associated with particular molecular activities and biological processes. Here, we present an updated and improved version, which builds on larger datasets of protein sequences and annotations, and uses updated component feature predictors as well as revised training procedures. FFPred 2.0 includes support vector regression models for the prediction of 442 Gene Ontology (GO) terms, which largely expand the coverage of the ontology and of the biological process category in particular. The GO term list mainly revolves around macromolecular interactions and their role in regulatory, signalling, developmental and metabolic processes. Benchmarking experiments on newly annotated proteins show that FFPred 2.0 provides more accurate functional assignments than its predecessor and the ProtFun server do; also, its assignments can complement information obtained using BLAST-based transfer of annotations, improving especially prediction in the biological process category. Furthermore, FFPred 2.0 can be used to annotate proteins belonging to several eukaryotic organisms with a limited decrease in prediction quality. We illustrate all these points through the use of both precision-recall plots and of the COGIC scores, which we recently proposed as an alternative numerical evaluation measure of function prediction accuracy.
DOI: 10.1006/jmbi.1999.3091
发表时间: 1999-09-17
影响因子: 5.6
作者:
Jones, DT
通讯作者: Jones, DT
DOI: 10.1016/s0022-2836(02)00379-0
发表时间: 2002-06-21
影响因子: 5.6
作者:
Jensen, LJ;Gupta, R;Brunak, S
通讯作者: Brunak, S
DOI: 10.1093/glycob/cwh151
发表时间: 2005-02-01
期刊: GLYCOBIOLOGY
影响因子: 4.3
作者:
Julenius, K;Molgaard, A;Brunak, S
通讯作者: Brunak, S
DOI: 10.1093/nar/gkr981
发表时间: 2012-01
影响因子: 14.9
作者:
UniProt Consortium
通讯作者: UniProt Consortium
DOI: 10.1007/s10994-007-5018-6
发表时间: 2007-10-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Lin, Hsuan-Tien;Lin, Chih-Jen;Weng, Ruby C.
通讯作者: Weng, Ruby C.