Exploiting ontology graph for predicting sparsely annotated gene function.

Exploiting ontology graph for predicting sparsely annotated gene function.
复制标题

利用本体图图来预测稀疏注释的基因功能。

DOI:
10.1093/bioinformatics/btv260
复制
发表时间:
2015-06-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Peng J
Peng J
中科院分区:
其他
文献类型:
--
作者:
Wang S;Cho H;Zhai C;Berger B;Peng J

文献摘要

被引文献

相似文献

动机:基于分子相互作用网络的系统预测基因(或蛋白质)功能已成为完善和增强现有注释目录的重要工具,如基因本体论(GO)数据库。然而,只有少数(<10)注释基因的功能标签构成了独特的挑战,因为任何独立考虑每个标签的预测算法都面临着信息匮乏,因此容易捕获数据中的不可概括模式,从而导致预测性能较差。存在用于函数预测的各种算法,但没有一种算法能够正确地解决稀疏注释函数的这种过适应问题,或者以可扩展到人类目录中的数万个函数的方式来这样做。结果:我们提出了一种新的函数预测算法clusDCA,它在相似的函数标签之间传递信息,以缓解稀疏标注函数的过拟合问题。我们的方法可扩展到具有大量注释的数据集。在酵母、小鼠和人类的交叉验证实验中,我们的方法在预测稀疏注释函数方面大大优于以前的最先进的函数预测算法,而不会牺牲对具有足够信息的标签的性能。此外,我们的方法仅基于本体图结构和与其他标签相关联的基因,可以准确地预测将被分配没有已知注释的功能标签的基因,这进一步表明我们的方法有效地利用了基因功能之间的相似性。可用性和实施:https://github.com/wangshenguiuc/clusDCA.联系方式:jianpeng@illinois.edu补充信息:补充数据可从BioInformation Online获得。
Motivation: Systematically predicting gene (or protein) function based on molecular interaction networks has become an important tool in refining and enhancing the existing annotation catalogs, such as the Gene Ontology (GO) database. However, functional labels with only a few (<10) annotated genes, which constitute about half of the GO terms in yeast, mouse and human, pose a unique challenge in that any prediction algorithm that independently considers each label faces a paucity of information and thus is prone to capture non-generalizable patterns in the data, resulting in poor predictive performance. There exist a variety of algorithms for function prediction, but none properly address this ‘overfitting’ issue of sparsely annotated functions, or do so in a manner scalable to tens of thousands of functions in the human catalog. Results: We propose a novel function prediction algorithm, clusDCA, which transfers information between similar functional labels to alleviate the overfitting problem for sparsely annotated functions. Our method is scalable to datasets with a large number of annotations. In a cross-validation experiment in yeast, mouse and human, our method greatly outperformed previous state-of-the-art function prediction algorithms in predicting sparsely annotated functions, without sacrificing the performance on labels with sufficient information. Furthermore, we show that our method can accurately predict genes that will be assigned a functional label that has no known annotations, based only on the ontology graph structure and genes associated with other labels, which further suggests that our method effectively utilizes the similarity between gene functions. Availability and implementation: https://github.com/wangshenguiuc/clusDCA. Contact: jianpeng@illinois.edu Supplementary information: Supplementary data are available at Bioinformatics online.