Distinguishing the species of biomedical named entities for term identification.

Distinguishing the species of biomedical named entities for term identification.
复制标题

DOI:
10.1186/1471-2105-9-s11-s6
复制
发表时间:
2008-11-19
期刊:
影响因子:
3
通讯作者:
Matthews, Michael
Matthews, Michael
中科院分区:
生物学4区
文献类型:
--
作者:
Wang, Xinglong;Matthews, Michael

文献摘要

被引文献

相似文献

术语识别是将文本中生物医学命名实体的模糊提及与唯一数据库标识符联系起来的任务。以前的术语识别工作主要集中在研究特定物种的文件。然而,全文文章通常描述跨多个物种的实体,在这种情况下,解决实体中模式生物的模糊性对于实现准确的术语识别至关重要。我们开发并比较了一些基于规则和基于机器学习的方法来解决生物医学命名实体中的物种模糊性,并证明了混合方法在黄金标准ITI-TXM语料库上测试时达到了71.7%的最佳整体准确度。通过利用混合标记器预测的物种信息,我们的基于规则的术语识别系统显着提高了11.6%。本文表明,在识别涉及多个模式生物的术语的背景下,整合一个准确的物种消歧系统可以显着提高术语识别系统的性能。
Term identification is the task of grounding ambiguous mentions of biomedical named entities in text to unique database identifiers. Previous work on term identification has focused on studying species-specific documents. However, full-length articles often describe entities across a number of species, in which case resolving the ambiguity of model organisms in entities is critical to achieving accurate term identification. We developed and compared a number of rule-based and machine-learning based approaches to resolving species ambiguity in mentions of biomedical named entities, and demonstrated that a hybrid method achieved the best overall accuracy at 71.7%, as tested on the gold-standard ITI-TXM corpora. By utilising the species information predicted by the hybrid tagger, our rule-based term identification system was improved significantly by up to 11.6%. This paper shows that, in the context of identifying terms involving multiple model organisms, integration of an accurate species disambiguation system can significantly improve the performance of term identification systems.