Improving subcellular localization prediction using text classification and the gene ontology

Improving subcellular localization prediction using text classification and the gene ontology
复制标题

DOI:
10.1093/bioinformatics/btn463
复制
发表时间:
2008-11-01
期刊:
影响因子:
5.8
通讯作者:
Lu, Paul
Lu, Paul
中科院分区:
生物学3区
文献类型:
--
作者:
Fyshe, Alona;Liu, Yifeng;Lu, Paul

文献摘要

被引文献

相似文献

动机:每种蛋白质都在细胞中的特定位置发挥其功能。这个亚细胞位置对于了解蛋白质的功能和促进其纯化是很重要的。现在有许多基于序列分析和来自同源基因的数据库信息来预测位置的计算技术。最近的一些技术使用来自生物摘要的文本:我们的目标是提高这种基于文本的技术的预测精度。我们确定了三种改进基于文本的预测的技术:歧义抽象去除规则,使用基因本体论(GO)中的同义词的机制,以及使用GO层次来泛化术语的机制。我们表明,这三种技术可以显着提高蛋白质亚细胞位置预测器的准确性,这些预测器使用从引用记录在Swiss-Prot中的PubMed摘要中提取的文本。
Motivation: Each protein performs its functions within some specific locations in a cell. This subcellular location is important for understanding protein function and for facilitating its purification. There are now many computational techniques for predicting location based on sequence analysis and database information from homologs. A few recent techniques use text from biological abstracts: our goal is to improve the prediction accuracy of such text-based techniques. We identify three techniques for improving text-based prediction: a rule for ambiguous abstract removal, a mechanism for using synonyms from the Gene Ontology (GO) and a mechanism for using the GO hierarchy to generalize terms. We show that these three techniques can significantly improve the accuracy of protein subcellular location predictors that use text extracted from PubMed abstracts whose references are recorded in Swiss-Prot.