Thesaurus-based disambiguation of gene symbols

Thesaurus-based disambiguation of gene symbols
复制标题

DOI:
10.1186/1471-2105-6-149
复制
发表时间:
2005-06-16
期刊:
影响因子:
3
通讯作者:
Kors, JA
Kors, JA
中科院分区:
生物学4区
文献类型:
--
作者:
Schijvenaars, BJA;Mons, B;Kors, JA

文献摘要

被引文献

相似文献

背景:生物学文献的大规模文本挖掘具有巨大的希望,可以将不同的信息和发现新知识联系起来。但是,基因符号的歧义是一个主要的瓶颈。回顾性:我们开发了一种基于词库的简单歧义算法,可以在很少的训练数据中运行。词库包括来自五个人类遗传数据库和网格的信息。人类基因符号的同源性问题的程度被证明是很大的(我们组合的词库中的33%的基因具有一个或多个模棱两可的符号),这不仅是因为一个符号可以指多个基因,还因为一个基因符号可以具有许多非基因含义。一组52,529个MEDLINE摘要,其中包含从OMIM中获得的690个模棱两可的人类基因符号。在测试集中,歧义算法的总体准确性高达92.7%。结论:人类基因符号的歧义是实质性的,这不仅是因为一个符号可能表示多个基因,而且特别是因为许多符号具有其他非基因含义。提出的歧义方法以高精度(包括重要的基因/不是基因决策)解决了我们的测试集中的大多数歧义。该算法是快速且可扩展的,在大规模文本挖掘应用中可以使基因符号歧义。
Background: Massive text mining of the biological literature holds great promise of relating disparate information and discovering new knowledge. However, disambiguation of gene symbols is a major bottleneck.Results: We developed a simple thesaurus-based disambiguation algorithm that can operate with very little training data. The thesaurus comprises the information from five human genetic databases and MeSH. The extent of the homonym problem for human gene symbols is shown to be substantial (33% of the genes in our combined thesaurus had one or more ambiguous symbols), not only because one symbol can refer to multiple genes, but also because a gene symbol can have many non-gene meanings. A test set of 52,529 Medline abstracts, containing 690 ambiguous human gene symbols taken from OMIM, was automatically generated. Overall accuracy of the disambiguation algorithm was up to 92.7% on the test set.Conclusion: The ambiguity of human gene symbols is substantial, not only because one symbol may denote multiple genes but particularly because many symbols have other, non-gene meanings. The proposed disambiguation approach resolves most ambiguities in our test set with high accuracy, including the important gene/not a gene decisions. The algorithm is fast and scalable, enabling gene-symbol disambiguation in massive text mining applications.