Protein function prediction by massive integration of evolutionary analyses and multiple data sources.

Protein function prediction by massive integration of evolutionary analyses and multiple data sources.
复制标题

DOI:
10.1186/1471-2105-14-s3-s1
复制
发表时间:
2013
期刊:
影响因子:
3
通讯作者:
Jones DT
Jones DT
中科院分区:
生物学4区
文献类型:
--
作者:
Cozzetto D;Buchan DW;Bryson K;Jones DT

文献摘要

被引文献

相似文献

准确的蛋白质功能注释是利用大量高通量下一代测序数据的严重瓶颈。保持数据库注释最新已成为一项重大的科学挑战,需要开发可靠的蛋白质功能自动预测器。CAFA实验提供了一个独特的机会,对许多不同的自动功能预测方法进行全面的“盲测”。我们报告了我们在这一挑战中使用的方法以及我们学到的经验教训。我们的方法集成到一个单一的框架,各种各样的生物信息源,包括序列,基因表达和蛋白质-蛋白质相互作用的数据,以及注释UniProt条目。该方法转移功能类别的基础上,从互补的同源性为基础的和功能为基础的分析结果。我们通过以概率的方式结合初始预测生成最终的分子功能和生物过程分配,该方式考虑了基因本体论的层次结构。我们提出了一种新的评分函数,称为组合图信息内容相似性(COGIC)得分预测的功能类别和基准数据的比较。我们证明,我们的综合方法提供了更大的范围和准确性的组件方法和天真的预测。与以前的研究一致,我们发现分子功能预测比生物过程分配更准确。总的来说,结果表明,外地有很大的改进余地。社区仍然需要投入大量的努力,使自动功能预测成为生命科学家工具箱中有用的常规组件。正如在其他领域已经看到的那样,全社区的盲测实验将在建立评估预测准确性的标准、促进进步和新思想以及最终记录进展方面发挥关键作用。
Accurate protein function annotation is a severe bottleneck when utilizing the deluge of high-throughput, next generation sequencing data. Keeping database annotations up-to-date has become a major scientific challenge that requires the development of reliable automatic predictors of protein function. The CAFA experiment provided a unique opportunity to undertake comprehensive 'blind testing' of many diverse approaches for automated function prediction. We report on the methodology we used for this challenge and on the lessons we learnt. Our method integrates into a single framework a wide variety of biological information sources, encompassing sequence, gene expression and protein-protein interaction data, as well as annotations in UniProt entries. The methodology transfers functional categories based on the results from complementary homology-based and feature-based analyses. We generated the final molecular function and biological process assignments by combining the initial predictions in a probabilistic manner, which takes into account the Gene Ontology hierarchical structure. We propose a novel scoring function called COmbined Graph-Information Content similarity (COGIC) score for the comparison of predicted functional categories and benchmark data. We demonstrate that our integrative approach provides increased scope and accuracy over both the component methods and the naïve predictors. In line with previous studies, we find that molecular function predictions are more accurate than biological process assignments. Overall, the results indicate that there is considerable room for improvement in the field. It still remains for the community to invest a great deal of effort to make automated function prediction a useful and routine component in the toolbox of life scientists. As already witnessed in other areas, community-wide blind testing experiments will be pivotal in establishing standards for the evaluation of prediction accuracy, in fostering advancements and new ideas, and ultimately in recording progress.