COMETA: A Corpus for Medical Entity Linking in the Social Media

COMETA: A Corpus for Medical Entity Linking in the Social Media
复制标题

DOI:
10.18653/v1/2020.emnlp-main.253
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Marco Basaldella;Fangyu Liu;Ehsan Shareghi;Nigel Collier
Marco Basaldella;Fangyu Liu;Ehsan Shareghi;Nigel Collier
中科院分区:
其他
文献类型:
--
作者:
Marco Basaldella;Fangyu Liu;Ehsan Shareghi;Nigel Collier

文献摘要

被引文献

相似文献

虽然在通用语言的实体链接(EL)方面取得了越来越大的进展,但现有的数据集无法解决外行语言中卫生术语的复杂性质。与此同时,在卫生领域,对能够理解公众声音的应用程序的需求也在不断增长。为了解决这个问题,我们引入了一个名为Cometa的新语料库,由来自Reddit Expert的20,000个英文生物医学实体提及组成-用到SNOMED CT的链接注释,SNOMED CT是一个广泛使用的医学知识图谱。我们的语料库满足了从规模和覆盖面到多样性和质量等令人满意的特性的组合,就我们所知,这是外地现有资源无法满足的。通过在20条EL基线上的基准实验,从字符串模型到神经模型,我们揭示了这些系统在两个具有挑战性的评估场景下对实体和概念执行复杂推理的能力。我们在Cometa上的实验结果表明,没有金弹存在,即使是最好的主流技术也有很大的性能差距需要填补,而最佳的解决方案依赖于结合不同的数据视图。
Whilst there has been growing progress in Entity Linking (EL) for general language, existing datasets fail to address the complex nature of health terminology in layman's language. Meanwhile, there is a growing need for applications that can understand the public's voice in the health domain. To address this we introduce a new corpus called COMETA, consisting of 20k English biomedical entity mentions from Reddit expert-annotated with links to SNOMED CT, a widely-used medical knowledge graph. Our corpus satisfies a combination of desirable properties, from scale and coverage to diversity and quality, that to the best of our knowledge has not been met by any of the existing resources in the field. Through benchmark experiments on 20 EL baselines from string- to neural-based models we shed light on the ability of these systems to perform complex inference on entities and concepts under 2 challenging evaluation scenarios. Our experimental results on COMETA illustrate that no golden bullet exists and even the best mainstream techniques still have a significant performance gap to fill, while the best solution relies on combining different views of data.