Assessment of disease named entity recognition on a corpus of annotated sentences.

Assessment of disease named entity recognition on a corpus of annotated sentences.
复制标题

DOI:
10.1186/1471-2105-9-s3-s3
复制
发表时间:
2008-04-11
期刊:
影响因子:
3
通讯作者:
Rebholz-Schuhmann D
Rebholz-Schuhmann D
中科院分区:
生物学4区
文献类型:
--
作者:
Jimeno A;Jimenez-Ruiz E;Lee V;Gaudan S;Berlanga R;Rebholz-Schuhmann D

文献摘要

被引文献

相似文献

近年来,从生物医学科学文献中识别语义类型一直集中在命名实体,如蛋白质和基因名称(PGNs)和基因本体术语(GO术语)。其他语义类型,如疾病,没有得到同样程度的关注。已经提出了不同的解决方案来识别科学文献中的疾病命名实体。虽然将术语与语言模式匹配遭受低召回(例如,Whatizit)其他解决方案利用形态句法特征来更好地覆盖术语可变性的全部范围(例如,MetaMap)。目前,由美国国家医学图书馆(NLM)提供的MetaMap是文献中UMLS(统一医学语言系统)概念注释的最先进解决方案。尽管如此,其性能尚未在注释语料库上进行评估。此外,到目前为止,几乎没有投入任何努力来生成一个注释数据集,该数据集将文本中的疾病实体与数据库、词库或本体中的疾病条目联系起来,并可以作为基准文本挖掘解决方案的黄金标准。作为我们研究工作的一部分,我们已经采取了语料库,已交付在过去的基础上UMLS元词库的基因与疾病的关联识别,我们已经重新处理和重新注释的语料库。我们从两位策展人那里收集了疾病实体的注释,分析了他们的分歧(kappa统计量为0.51),并组成了一个单独的注释语料库供公众使用。此后,疾病命名实体识别的三种解决方案,包括MetaMap已被应用到语料库自动注释它与UMLS元词库的概念。所得到的注释已经过基准测试以比较它们的性能。注释语料库可在以下网址公开获取: 并且可以作为其他系统的基准。此外,我们发现,字典查找已经提供了有竞争力的结果,表明疾病术语的使用是高度标准化的整个术语和文献。MetaMap以召回率不足为代价生成精确的结果,而我们的统计方法以较低的准确率获得更好的召回率。通过组合三种方法中的至少两种,可以在精度方面获得更好的结果,但这种方法再次降低了召回率。总而言之,我们的分析可以更好地了解文献中疾病注释的复杂性。MetaMap和基于字典的方法可通过Whatizit web服务基础设施获得(Rebholz-Schuhmann D,Arregui M,Gaudan S,Kirsch H,Jimeno A:Text processing through Web services:Calling Whatizit. Bioinformatics 2008,24:296-298)。
In recent years, the recognition of semantic types from the biomedical scientific literature has been focused on named entities like protein and gene names (PGNs) and gene ontology terms (GO terms). Other semantic types like diseases have not received the same level of attention. Different solutions have been proposed to identify disease named entities in the scientific literature. While matching the terminology with language patterns suffers from low recall (e.g., Whatizit) other solutions make use of morpho-syntactic features to better cover the full scope of terminological variability (e.g., MetaMap). Currently, MetaMap that is provided from the National Library of Medicine (NLM) is the state of the art solution for the annotation of concepts from UMLS (Unified Medical Language System) in the literature. Nonetheless, its performance has not yet been assessed on an annotated corpus. In addition, little effort has been invested so far to generate an annotated dataset that links disease entities in text to disease entries in a database, thesaurus or ontology and that could serve as a gold standard to benchmark text mining solutions. As part of our research work, we have taken a corpus that has been delivered in the past for the identification of associations of genes to diseases based on the UMLS Metathesaurus and we have reprocessed and re-annotated the corpus. We have gathered annotations for disease entities from two curators, analyzed their disagreement (0.51 in the kappa-statistic) and composed a single annotated corpus for public use. Thereafter, three solutions for disease named entity recognition including MetaMap have been applied to the corpus to automatically annotate it with UMLS Metathesaurus concepts. The resulting annotations have been benchmarked to compare their performance. The annotated corpus is publicly available at and can serve as a benchmark to other systems. In addition, we found that dictionary look-up already provides competitive results indicating that the use of disease terminology is highly standardized throughout the terminologies and the literature. MetaMap generates precise results at the expense of insufficient recall while our statistical method obtains better recall at a lower precision rate. Even better results in terms of precision are achieved by combining at least two of the three methods leading, but this approach again lowers recall. Altogether, our analysis gives a better understanding of the complexity of disease annotations in the literature. MetaMap and the dictionary based approach are available through the Whatizit web service infrastructure (Rebholz-Schuhmann D, Arregui M, Gaudan S, Kirsch H, Jimeno A: Text processing through Web services: Calling Whatizit. Bioinformatics 2008, 24:296-298).