Prediction of Metabolic Pathway Involvement in Prokaryotic UniProtKB Data by Association Rule Mining.

Prediction of Metabolic Pathway Involvement in Prokaryotic UniProtKB Data by Association Rule Mining.
复制标题

DOI:
10.1371/journal.pone.0158896
复制
发表时间:
2016
期刊:
影响因子:
3.7
通讯作者:
Solovyev V
Solovyev V
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Boudellioua I;Saidi R;Hoehndorf R;Martin MJ;Solovyev V

文献摘要

被引文献

相似文献

已知蛋白质与其功能之间的差距不断扩大,这鼓励了自动推断注释的方法的发展。蛋白质的自动功能注释有望满足最大化注释覆盖率的矛盾要求,同时最大限度地减少错误的功能分配。这种权衡给设计智能系统以解决自动蛋白质注释问题带来了巨大挑战。在这项工作中,我们提出了一个系统,利用规则挖掘技术来预测原核生物的代谢途径。所得到的知识表示将通路参与分配给UniProtKB条目的预测模型。我们进行了评估研究,我们的系统性能使用交叉验证技术。我们发现,它取得了非常有前途的结果,在途径鉴定与F1的措施为0.982和AUC为0.987。我们的预测模型,然后成功地应用于620万UniProtKB/TrEMBL参考蛋白质组条目的原核生物。结果,涵盖了663,724个条目,其中436,510个条目缺乏任何以前的路径注释。
The widening gap between known proteins and their functions has encouraged the development of methods to automatically infer annotations. Automatic functional annotation of proteins is expected to meet the conflicting requirements of maximizing annotation coverage, while minimizing erroneous functional assignments. This trade-off imposes a great challenge in designing intelligent systems to tackle the problem of automatic protein annotation. In this work, we present a system that utilizes rule mining techniques to predict metabolic pathways in prokaryotes. The resulting knowledge represents predictive models that assign pathway involvement to UniProtKB entries. We carried out an evaluation study of our system performance using cross-validation technique. We found that it achieved very promising results in pathway identification with an F1-measure of 0.982 and an AUC of 0.987. Our prediction models were then successfully applied to 6.2 million UniProtKB/TrEMBL reference proteome entries of prokaryotes. As a result, 663,724 entries were covered, where 436,510 of them lacked any previous pathway annotations.