A reliability studs for evaluating information extraction from radiology reports

A reliability studs for evaluating information extraction from radiology reports
复制标题

DOI:
10.1136/jamia.1999.0060143
复制
发表时间:
1999-03-01
影响因子:
6.4
通讯作者:
Heitjan, DF
Heitjan, DF
中科院分区:
管理学2区
文献类型:
--
作者:
Hripcsak, G;Kuperman, GJ;Heitjan, DF

文献摘要

被引文献

相似文献

目的:评估信息提取任务参考标准的可靠性。设置:来自两个地点和两个专科的24名医生评分员通过阅读胸部X线片报告来判断是否存在临床情况。方法:估计方差成分、概化(可靠性)系数和生成可靠参考标准所需的专家评分员数量。结果:每个评分员的平均可靠性为0.80(95%CI,0.79-0.81)。9种单独情况的可靠性从0.67到0.97不等,中心线存在和气胸最可靠,胸腔积液(不包括充血性心力衰竭)和肺炎最不可靠。需要一到两个评分员才能达到0.70的可靠性,平均需要六个评分员才能达到0.95的可靠性。对于更复杂的任务,这比之前公布的每个评分者0.19的可靠性要可靠得多。结论:在这些评估中,医生评分者能够基于文本报告非常可靠地判断临床条件的存在。一旦特定评分者的可靠性得到确认,该评分者就有可能创建一个足够可靠的参考标准,以评估系统上的综合衡量标准。需要六名评分员才能制定一个参考标准,足以在个案基础上评估一个系统。这些结果应该有助于评估者为自然语言处理器和其他基于知识的系统设计未来的信息提取研究。
Goal: To assess the reliability of a reference standard for an information extraction task.Setting: Twenty-four physician raters from two sites and two specialties judged whether clinical conditions were present based on reading chest radiograph reports.Methods: Variance components, generalizability (reliability) coefficients, and the number of expert raters needed to generate a reliable reference standard were estimated.Results: Per-rater reliability averaged across conditions was 0.80 (95% CI, 0.79-0.81). Reliability for the nine individual conditions varied from 0.67 to 0.97, with central line presence and pneumothorax the most reliable, and pleural effusion (excluding CHF) and pneumonia the least reliable. One to two raters were needed to achieve a reliability of 0.70, and six raters, on average, were required to achieve a reliability of 0.95. This was far more reliable than a previously published per-rater reliability of 0.19 for a more complex task. Differences between sites were attributable to changes to the condition definitions.Conclusion: In these evaluations, physician raters were able to judge very reliably the presence of clinical conditions based on text reports. Once the reliability of a specific rater is confirmed, it would be possible for that rater to create a reference standard reliable enough to assess aggregate measures on a system. Six raters would be needed to create a reference standard sufficient to assess a system on a case-by-case basis. These results should help evaluators design future information extraction studies for natural language processors and other knowledge-based systems.