Automatically annotating documents with normalized gene lists.

Automatically annotating documents with normalized gene lists.
复制标题

DOI:
10.1186/1471-2105-6-s1-s13
复制
发表时间:
2005
期刊:
影响因子:
3
通讯作者:
Pereira F
Pereira F
中科院分区:
生物学4区
文献类型:
--
作者:
Crim J;McDonald R;Pereira F

文献摘要

被引文献

相似文献

文档基因规范化是为文档中提到的基因创建唯一标识符列表的问题。自动化这个过程在信息提取和数据库管理系统中有许多潜在的应用。在这里,我们提出了两个不同的解决方案,这个问题。第一种方法主要基于标准模式匹配和信息提取技术。第二种更新颖的解决方案使用统计分类器从已知基因同义词列表中识别有效的基因匹配。我们比较了两个系统的结果,分析了它们的优点,并认为基于分类的系统是更好的,包括性能,简单性和鲁棒性的许多原因。我们最好的系统可以达到74%-92%的精确度和召回率,具体取决于生物体。
Document gene normalization is the problem of creating a list of unique identifiers for genes that are mentioned within a document. Automating this process has many potential applications in both information extraction and database curation systems. Here we present two separate solutions to this problem. The first is primarily based on standard pattern matching and information extraction techniques. The second and more novel solution uses a statistical classifier to recognize valid gene matches from a list of known gene synonyms. We compare the results of the two systems, analyze their merits and argue that the classification based system is preferable for many reasons including performance, simplicity and robustness. Our best systems attain a balanced precision and recall in the range of 74%–92%, depending on the organism.