Automatic concept recognition using the human phenotype ontology reference and test suite corpora.

Automatic concept recognition using the human phenotype ontology reference and test suite corpora.
复制标题

DOI:
10.1093/database/bav005
复制
发表时间:
2015-01-01
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Robinson, Peter N
Robinson, Peter N
中科院分区:
其他
文献类型:
--
作者:
Groza, Tudor;Kohler, Sebastian;Robinson, Peter N

文献摘要

被引文献

相似文献

概念识别工具依赖于文本语料库的可用性来评估其性能并识别需要改进的领域。通常,语料库是为了特定目的而开发的,例如基因名称识别。基因和蛋白质名称识别是生物医学文本挖掘的长期目标,因此存在许多不同的语料库。然而,表型直到最近才成为专门概念识别系统感兴趣的实体,并且几乎没有任何注释文本可用于性能测试和培训。在这里,我们提出了一个独特的语料库,捕获来自 228 个摘要的文本跨度,这些摘要用人类表型本体论 (HPO) 概念手动注释,并由三位策展人协调,可以用作人类表型的自由文本注释的参考标准。此外,我们开发了一个用于标准化概念识别错误分析的测试套件,包含与 2164 个 HPO 概念相对应的 32 种不同类型的测试用例。最后,对三个已建立的表型概念识别器(NCBO Annotator、OBO Annotator 和 Bio-LarK CR)进行了综合评估,并根据文本语料库和测试套件报告结果。黄金标准和测试套件语料库可从 http://bio-lark.org/hpo_res.html 获取。数据库网址:http://bio-lark.org/hpo_res.html。
Concept recognition tools rely on the availability of textual corpora to assess their performance and enable the identification of areas for improvement. Typically, corpora are developed for specific purposes, such as gene name recognition. Gene and protein name identification are longstanding goals of biomedical text mining, and therefore a number of different corpora exist. However, phenotypes only recently became an entity of interest for specialized concept recognition systems, and hardly any annotated text is available for performance testing and training. Here, we present a unique corpus, capturing text spans from 228 abstracts manually annotated with Human Phenotype Ontology (HPO) concepts and harmonized by three curators, which can be used as a reference standard for free text annotation of human phenotypes. Furthermore, we developed a test suite for standardized concept recognition error analysis, incorporating 32 different types of test cases corresponding to 2164 HPO concepts. Finally, three established phenotype concept recognizers (NCBO Annotator, OBO Annotator and Bio-LarK CR) were comprehensively evaluated, and results are reported against both the text corpus and the test suites. The gold standard and test suites corpora are available from http://bio-lark.org/hpo_res.html. Database URL: http://bio-lark.org/hpo_res.html.