Assessing protein similarity with Gene Ontology and its use in subnuclear localization prediction

Assessing protein similarity with Gene Ontology and its use in subnuclear localization prediction
复制标题

DOI:
10.1186/1471-2105-7-491
复制
发表时间:
2006-11-07
期刊:
影响因子:
3
通讯作者:
Dai, Yang
Dai, Yang
中科院分区:
生物学4区
文献类型:
--
作者:
Lei, Zhengdeng;Dai, Yang

文献摘要

被引文献

相似文献

背景:各种基因组测序项目的完成积累了大量的基因序列信息。这就需要一种大规模的计算方法来从序列中预测蛋白质的定位。蛋白质的定位可以提供关于它的分子功能以及它参与的生物途径的有价值的信息。蛋白质在亚核水平的定位预测是一项具有挑战性的任务。在我们之前的工作中,我们提出了一个基于支持向量机的系统,该系统利用蛋白质序列信息来进行预测。在这项工作中,我们使用基因本体论(GO)来评估蛋白质的相似性,然后通过增加最近邻分类器模块来提高系统的性能。结果:使用来自核蛋白数据库(NPD)的6个定位的蛋白质集,本文提出的新系统的性能与我们以前的系统进行了比较。在留一法交叉验证中,单一定位蛋白的总体MCC(准确率)从0.284(50.0%)提高到0.519(66.5%);对于一组独立的多定位蛋白质,总的MCC(准确率)从0.420(65.2%)提高到0.541(65.2%)。新系统可在http://array.bioengr.uic.edu/subnuclear.htm上获得。结论:基于GO术语的不同相似性度量对蛋白质对相似性的不同定义可能在很大程度上影响蛋白质亚核定位的预测。使用两个蛋白质的匹配GO术语对的相似性分数之和作为相似性定义产生了最好的预测结果。通过将基因本体论与序列信息相结合,在预测蛋白质亚核定位方面取得了实质性的改进。
Background: The accomplishment of the various genome sequencing projects resulted in accumulation of massive amount of gene sequence information. This calls for a large-scale computational method for predicting protein localization from sequence. The protein localization can provide valuable information about its molecular function, as well as the biological pathway in which it participates. The prediction of localization of a protein at subnuclear level is a challenging task. In our previous work we proposed an SVM-based system using protein sequence information for this prediction task. In this work, we assess protein similarity with Gene Ontology (GO) and then improve the performance of the system by adding a module of nearest neighbor classifier using a similarity measure derived from the GO annotation terms for protein sequences.Results: The performance of the new system proposed here was compared with our previous system using a set of proteins resided within 6 localizations collected from the Nuclear Protein Database (NPD). The overall MCC (accuracy) is elevated from 0.284 (50.0%) to 0.519 (66.5%) for single-localization proteins in leave-one-out cross-validation; and from 0.420 (65.2%) to 0.541 (65.2%) for an independent set of multi-localization proteins. The new system is available at http://array.bioengr.uic.edu/ subnuclear.htm.Conclusion: The prediction of protein subnuclear localizations can be largely influenced by various definitions of similarity for a pair of proteins based on different similarity measures of GO terms. Using the sum of similarity scores over the matched GO term pairs for two proteins as the similarity definition produced the best predictive outcome. Substantial improvement in predicting protein subnuclear localizations has been achieved by combining Gene Ontology with sequence information.