Community annotation experiment for ground truth generation for the i2b2 medication challenge

Community annotation experiment for ground truth generation for the i2b2 medication challenge
复制标题

DOI:
10.1136/jamia.2010.004200
复制
发表时间:
2010-09-01
影响因子:
6.4
通讯作者:
Cadag, Eithon
Cadag, Eithon
中科院分区:
管理学2区
文献类型:
--
作者:
Uzuner, Oezlem;Solti, Imre;Cadag, Eithon

文献摘要

被引文献

相似文献

目的 在第三届 i2b2 临床记录自然语言处理挑战研讨会的背景下,作者(也称为“i2b2 药物挑战团队”或简称“i2b2 团队”)组织了一次社区注释实验。 设计 对于该实验,作者发布了注释指南和一小组带注释的出院摘要。他们要求第三次 i2b2 研讨会的参与者每人注释 10 份出院小结;每个出院摘要均由来自两个不同团队的两名注释者进行注释,第三个团队的第三个注释者解决了分歧。 测量为了评估由此产生的注释的可靠性,作者测量了社区注释者间的一致性,并将其与当社区和专家注释者基于汇集的系统输出生成地面事实时专家注释者的注释者间一致性进行比较。为此,该池由每条记录的三个最密集的自动注释组成。当专家在不使用池的情况下注释原始记录时,作者还将社区注释者间协议与专家注释者间协议进行了比较。最后,他们通过将社区基本事实与专家基本事实进行比较来衡量社区基本事实的质量。 结果和结论 作者发现,无论专家是否从池中进行注释,社区注释者都实现了与专家注释者相当的注释者间一致性。此外,社区生成的地面事实相对于专家的地面事实获得了 0.90 以上的 F 度量,表明即使在复杂且特定领域的注释任务中,社区作为高质量地面事实来源的价值也是如此。
Objective Within the context of the Third i2b2 Workshop on Natural Language Processing Challenges for Clinical Records, the authors (also referred to as 'the i2b2 medication challenge team' or 'the i2b2 team' for short) organized a community annotation experiment.Design For this experiment, the authors released annotation guidelines and a small set of annotated discharge summaries. They asked the participants of the Third i2b2 Workshop to annotate 10 discharge summaries per person; each discharge summary was annotated by two annotators from two different teams, and a third annotator from a third team resolved disagreements.Measurements In order to evaluate the reliability of the annotations thus produced, the authors measured community inter-annotator agreement and compared it with the inter-annotator agreement of expert annotators when both the community and the expert annotators generated ground truth based on pooled system outputs. For this purpose, the pool consisted of the three most densely populated automatic annotations of each record. The authors also compared the community inter-annotator agreement with expert inter-annotator agreement when the experts annotated raw records without using the pool. Finally, they measured the quality of the community ground truth by comparing it with the expert ground truth.Results and conclusions The authors found that the community annotators achieved comparable inter-annotator agreement to expert annotators, regardless of whether the experts annotated from the pool. Furthermore, the ground truth generated by the community obtained F-measures above 0.90 against the ground truth of the experts, indicating the value of the community as a source of high-quality ground truth even on intricate and domain-specific annotation tasks.