OGER plus plus : hybrid multi-type entity recognition

OGER plus plus : hybrid multi-type entity recognition
复制标题

DOI:
10.1186/s13321-018-0326-3
复制
发表时间:
2019-01-21
影响因子:
8.6
通讯作者:
Rinaldi, Fabio
Rinaldi, Fabio
中科院分区:
化学2区
文献类型:
--
作者:
Furrer, Lenz;Jancso, Anna;Rinaldi, Fabio

文献摘要

被引文献

相似文献

背景:我们提出了一个文本挖掘工具,用于识别科学文献中的生物医学实体。oger++是一个用于命名实体识别和概念识别(链接)的混合系统,它结合了基于字典的注释器和基于语料库的消歧组件。注释器使用有效的查找策略和匹配拼写变体的规范化方法相结合。消歧分类器作为前馈神经网络实现,作为前一步的后过滤器。结果:我们从处理速度和标注质量两方面对系统进行了评价。在速度基准测试中,oger++ web服务每秒处理9.7个摘要或0.9个全文文档。在CRAFT语料库上,命名实体识别和概念识别的F1分别达到了71.4%和56.7%。结论:结合基于知识和数据驱动的组件,可以创建一个具有竞争力性能的生物医学文本挖掘系统。
Background: We present a text-mining tool for recognizing biomedical entities in scientific literature. OGER++ is a hybrid system for named entity recognition and concept recognition (linking), which combines a dictionary-based annotator with a corpus-based disambiguation component. The annotator uses an efficient look-up strategy combined with a normalization method for matching spelling variants. The disambiguation classifier is implemented as a feed-forward neural network which acts as a postfilter to the previous step.Results: We evaluated the system in terms of processing speed and annotation quality. In the speed benchmarks, the OGER++ web service processes 9.7 abstracts or 0.9 full-text documents per second. On the CRAFT corpus, we achieved 71.4% and 56.7% F1 for named entity recognition and concept recognition, respectively.Conclusions: Combining knowledge-based and data-driven components allows creating a system with competitive performance in biomedical text mining.