Disambiguating the species of biomedical named entities using natural language parsers.

Disambiguating the species of biomedical named entities using natural language parsers.
复制标题

DOI:
10.1093/bioinformatics/btq002
复制
发表时间:
2010-03-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Ananiadou S
Ananiadou S
中科院分区:
其他
文献类型:
--
作者:
Wang X;Tsujii J;Ananiadou S

文献摘要

参考文献

被引文献

相似文献

动机:文本挖掘技术已被证明可以减少组织隐藏在文献中的大量信息所涉及的繁重工作。文本挖掘中的一个挑战是将模糊的单词形式与明确的生物概念联系起来。本文报告了一项全面的研究,解决生物医学命名实体的模式生物方面提到的歧义,并提出了一系列的方法,重点是利用自然语言解析器的方法。结果如下:我们建立了一个语料库的生物消歧,其中每一个出现的蛋白质/基因实体手动标记的物种ID,并评估了一些方法上it. Promising的结果是通过训练机器学习模型的句法分析树,然后使用它来决定一个实体是否属于模型生物表示的相邻物种指示词(如酵母)。基于解析器的方法也进行了比较监督分类方法和结果表明,前者是一个更有利的选择时,域可移植性的关注。通过结合句法特征和监督分类的优势,获得了最佳的整体性能。可用性:语料库和演示可在http://www.nactem.ac.uk/deca_details/start.cgi上获得,该软件可作为U-Compare组件免费获得(Kano等人,):NaCTeM物种词检测器和NaCTeM物种消歧器。U-Compare可在http://-compare.org/上xinglong.wang @ manchester.ac.uk
Motivation: Text mining technologies have been shown to reduce the laborious work involved in organizing the vast amount of information hidden in the literature. One challenge in text mining is linking ambiguous word forms to unambiguous biological concepts. This article reports on a comprehensive study on resolving the ambiguity in mentions of biomedical named entities with respect to model organisms and presents an array of approaches, with focus on methods utilizing natural language parsers. Results: We build a corpus for organism disambiguation where every occurrence of protein/gene entity is manually tagged with a species ID, and evaluate a number of methods on it. Promising results are obtained by training a machine learning model on syntactic parse trees, which is then used to decide whether an entity belongs to the model organism denoted by a neighbouring species-indicating word (e.g. yeast). The parser-based approaches are also compared with a supervised classification method and results indicate that the former are a more favorable choice when domain portability is of concern. The best overall performance is obtained by combining the strengths of syntactic features and supervised classification. Availability: The corpus and demo are available at http://www.nactem.ac.uk/deca_details/start.cgi, and the software is freely available as U-Compare components (Kano et al.,): NaCTeM Species Word Detector and NaCTeM Species Disambiguator. U-Compare is available at http://-compare.org/ Contact: xinglong.wang@manchester.ac.uk
DOI: 10.1186/1471-2105-9-s11-s6
发表时间: 2008-11-19
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Wang, Xinglong;Matthews, Michael
通讯作者: Matthews, Michael
评估生物学的文本挖掘系统:第二次生物综合社区挑战的概述。
DOI: 10.1186/gb-2008-9-s2-s1
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Krallinger M;Morgan A;Smith L;Leitner F;Tanabe L;Wilbur J;Hirschman L;Valencia A
通讯作者: Valencia A
DOI: 10.1093/bioinformatics/bth496
发表时间: 2005-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Chen, LF;Liu, HF;Friedman, C
通讯作者: Friedman, C
DOI: 10.1093/bioinformatics/bti475
发表时间: 2005-07-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Settles, B
通讯作者: Settles, B
DOI: 10.1093/bioinformatics/btp289
发表时间: 2009-08-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Kano Y;Baumgartner WA Jr;McCrohon L;Ananiadou S;Cohen KB;Hunter L;Tsujii J
通讯作者: Tsujii J