Evaluating the state-of-the-art in automatic de-identification

Evaluating the state-of-the-art in automatic de-identification
复制标题

DOI:
10.1197/jamia.m2444
复制
发表时间:
2007-09-01
影响因子:
6.4
通讯作者:
Szolovits, Peter
Szolovits, Peter
中科院分区:
管理学2区
文献类型:
--
作者:
Uzuner, Oezlem;Luo, Yuan;Szolovits, Peter

文献摘要

被引文献

相似文献

为了促进和调查自动去身份识别的研究,作为 i2b2(将生物学整合到床边的信息学)项目的一部分,作者组织了一场自然语言处理 (NLP) 挑战赛,主题是从医疗出院记录中自动删除私人健康信息 (PHI)。本手稿概述了这一去识别化挑战,描述了数据和注释过程,解释了评估指标,讨论了应对挑战的系统的性质,分析了收到的系统运行的结果,并确定了未来研究的方向。去识别化挑战数据包括从合作伙伴医疗保健系统中提取的出院摘要。作者通过用合成替代物替换真实的 PHI 来准备这些数据以应对挑战。为了将挑战集中在非基于字典的去识别方法上,我们使用词汇表外的 PHI 替代项(即虚构的名称)来丰富数据。数据还包括一些与医学非 PHI 术语不明确的 PHI 替代项。共有七支队伍参加了此次挑战。每个团队最多提交 3 次系统运行,总共 16 次提交。作者使用精确度、召回率和 F 度量来根据提交的系统在真实情况下的标记级和实例级性能来评估其运行情况。性能最佳的系统在所有 PHI 类别的 F 测量中得分均超过 98%。大多数词汇表之外的 PHI 都可以被准确识别。然而,事实证明,识别不明确的 PHI 具有挑战性。系统在测试数据集上的表现令人鼓舞。这些系统的未来评估将涉及来自更多异构来源的更大数据集。
To facilitate and survey studies in automatic de-identification, as a part of the i2b2 (Informatics for Integrating Biology to the Bedside) project, authors organized a Natural Language Processing (NLP) challenge on automatically removing private health information (PHI) from medical discharge records. This manuscript provides an overview of this de-identification challenge, describes the data and the annotation process, explains the evaluation metrics, discusses the nature of the systems that addressed the challenge, analyzes the results of received system runs, and identifies directions for future research. The de-indentification challenge data consisted of discharge summaries drawn from the Partners Healthcare system. Authors prepared this data for the challenge by replacing authentic PHI with synthesized surrogates. To focus the challenge on non-dictionary-based de-identification methods, the data was enriched with out-of-vocabulary PHI surrogates, i.e., made up names. The data also included some PHI surrogates that were ambiguous with medical non-PHI terms. A total of seven teams participated in the challenge. Each team submitted up to three system runs, for a total of sixteen submissions. The authors used precision, recall, and F-measure to evaluate the submitted system runs based on their token-level and instance-level performance on the ground truth. The systems with the best performance scored above 98% in F-measure for all categories of PHI. Most out-of-vocabulary PHI could be identified accurately. However, identifying ambiguous PHI proved challenging. The performance of systems on the test data set is encouraging. Future evaluations of these systems will involve larger data sets from more heterogeneous sources.