Deriving an English Biomedical Silver Standard Corpus for CLEF-ER

Deriving an English Biomedical Silver Standard Corpus for CLEF-ER
复制标题

为 CLEF-ER 衍生英语生物医学银标准语料库

DOI:
10.5167/uzh-87213
复制
发表时间:
2013
影响因子:
1.9
通讯作者:
S. Clematide
S. Clematide
中科院分区:
工程技术4区
文献类型:
--
作者:
Ian Lewin;S. Clematide

文献摘要

被引文献

相似文献

我们描述了用于构建英语银标准注释的自动协调方法,该注释作为多语言CLEF-ER命名实体识别挑战的数据源。自动银标准的使用旨在消除对昂贵且耗时的专家注释的需要。来自项目合作伙伴的6个不同注释的统一最终投票门槛为3,保留了所有可用概念质心的45%。平均有19% (SD 14%)的原始注释被删除。进入银标准语料库的97.8%的伙伴注释具有与其协调表示完全相同的边界。
We describe the automatic harmonization method used for building the English Silver Standard annotation supplied as a data source for the multilingual CLEF-ER named entity recognition challenge. The use of an automatic Silver Standard is designed to remove the need for a costly and time-consuming expert annotation. The final voting threshold of 3 for the harmonization of 6 different annotations from the project partners kept 45% of all available concept centroids. On average, 19% (SD 14%) of the original annotations are removed. 97.8% of the partner annotations that go into the Silver Standard Corpus have exactly the same boundaries as their harmonized representations.