Clinical Information Extraction at the CLEF eHealth Evaluation lab 2016

Clinical Information Extraction at the CLEF eHealth Evaluation lab 2016
复制标题

DOI:
--
复制
发表时间:
2016-09
期刊:
CEUR workshop proceedings
影响因子:
--
通讯作者:
Aurélie Névéol;K. Cohen;Cyril Grouin;Thierry Hamon;T. Lavergne;Liadh Kelly;Lorraine Goeuriot;G. Rey;Aude Robert;Xavier Tannier;Pierre Zweigenbaum
Aurélie Névéol;K. Cohen;Cyril Grouin;Thierry Hamon;T. Lavergne;Liadh Kelly;Lorraine Goeuriot;G. Rey;Aude Robert;Xavier Tannier;Pierre Zweigenbaum
中科院分区:
其他
文献类型:
--
作者:
Aurélie Névéol;K. Cohen;Cyril Grouin;Thierry Hamon;T. Lavergne;Liadh Kelly;Lorraine Goeuriot;G. Rey;Aude Robert;Xavier Tannier;Pierre Zweigenbaum

文献摘要

被引文献

相似文献

本文报告了 2016 年 CLEF eHealth 评估实验室的任务 2,该任务扩展了 ShaARe/CLEF eHealth 评估实验室之前的信息提取任务。该任务继续进行法语叙述中的命名实体识别和标准化,如 CLEF eHealth 2015 中所提供的。命名实体识别涉及十种类型的实体,包括根据统一医学语言系统® (UMLS®) 中的语义组定义的疾病,该系统也用于实体的标准化。此外,我们在法国死亡证明中引入了一项大规模分类任务,其中包括提取国际疾病分类第十次修订版 (ICD10) 中编码的死亡原因。使用精确度、召回率和 F 测量,根据 MEDLINE 索引的 832 篇科学文章标题、欧洲药品管理局 (EMEA) 出版的 4 篇药物专着和 27,850 份死亡证明的盲参考标准对参与者系统进行评估。共有七个团队参与,其中实体识别和规范化任务有五个团队,死亡证明编码任务有五个团队。三个团队将他们的系统提交给我们新提供的再现性轨道。对于实体识别,在 EMEA 语料库上实现了最高性能,普通实体识别的总体 F 度量为 0.702,标准化实体识别的总体 F 度量为 0.529。对于实体标准化,MEDLINE 语料库上实现了最高性能,总体 F 度量为 0.552。对于死亡证明编码,最高性能为 0.848 F 测量。
This paper reports on Task 2 of the 2016 CLEF eHealth evaluation lab which extended the previous information extraction tasks of ShARe/CLEF eHealth evaluation labs. The task continued with named entity recognition and normalization in French narratives, as offered in CLEF eHealth 2015. Named entity recognition involved ten types of entities including disorders that were defined according to Semantic Groups in the Unified Medical Language System® (UMLS®), which was also used for normalizing the entities. In addition, we introduced a large-scale classification task in French death certificates, which consisted of extracting causes of death as coded in the International Classification of Diseases, tenth revision (ICD10). Participant systems were evaluated against a blind reference standard of 832 titles of scientific articles indexed in MEDLINE, 4 drug monographs published by the European Medicines Agency (EMEA) and 27,850 death certificates using Precision, Recall and F-measure. In total, seven teams participated, including five in the entity recognition and normalization task, and five in the death certificate coding task. Three teams submitted their systems to our newly offered reproducibility track. For entity recognition, the highest performance was achieved on the EMEA corpus, with an overall F-measure of 0.702 for plain entities recognition and 0.529 for normalized entity recognition. For entity normalization, the highest performance was achieved on the MEDLINE corpus, with an overall F-measure of 0.552. For death certificate coding, the highest performance was 0.848 F-measure.