A de-identifier for medical discharge summaries

A de-identifier for medical discharge summaries
复制标题

DOI:
10.1016/j.artmed.2007.10.001
复制
发表时间:
2008-01-01
影响因子:
7.5
通讯作者:
Szovits, Peter
Szovits, Peter
中科院分区:
工程技术1区
文献类型:
--
作者:
Uzuner, Oezlem;Sibanda, Tawanda C.;Szovits, Peter

文献摘要

被引文献

相似文献

目的:临床记录包含重要的医学信息,可用于各学科的研究人员。但是,这些记录还包含个人健康信息(PHI),这些信息的存在限制了这些记录在医院之外的使用。去标识化的目标是从临床记录中删除所有PHI。这是一项具有挑战性的任务,因为许多记录包含外国和拼写错误的PHI;它们还包含与非PHI不明确的PHI。临床记录的语言特征使这些并发症更加复杂。例如,本文所研究的医学出院摘要,其特点是话语碎片化、不完整和特定领域语言;它们不能被为Lay语言设计的工具完全处理。方法和结果:在本文中,我们证明了我们可以使用基于支持向量机和局部上下文的去标识符Stat De-id去标识医疗出院摘要(在PHI上F-measure = 97%)。我们对本地上下文的表示有助于去识别,即使PHI包含词汇外的单词,即使PHI与同一语料库中的非PHI模棱两可。Stat De-id与基于规则的方法的比较表明,本地上下文对去识别的贡献大于字典与手工定制的启发式相结合(F-measure = 85%)。与两个著名的命名实体识别(NER)系统,SNoW (F-measure = 94%)和IdentiFinder (F-measure = 36%)在五个代表性语料库上的比较表明,当文档语言支离破碎时,一个相对彻底地表示本地上下文的系统比将(相对简单的)本地上下文与全局上下文结合起来的系统更有效地去标识符。与条件随机场去标识符(CRFD)进行比较,该方法除了利用Stat De-id的本地上下文外,还利用了全局上下文,证实了这一发现(F-measure=88%),并确立了加强本地上下文的表示可能比用全局上下文补充本地上下文更有利于去标识。(C) 2007 Elsevier B.V.版权所有
Objective: Clinical records contain significant medical information that can be useful to researchers in various disciplines. However, these records also contain personal health information (PHI) whose presence limits the use of the records outside of hospitals.The goal of de-identification is to remove all PHI from clinical records. This is a challenging task because many records contain foreign and misspelled PHI; they also contain PHI that are ambiguous with non-PHI. These complications are compounded by the linguistic characteristics of clinical records. For example, medical discharge summaries, which are studied in this paper, are characterized by fragmented, incomplete utterances and domain-specific language; they cannot be fully processed by tools designed for Lay language.Methods and results: In this paper, we show that we can de-identify medical discharge summaries using a de-identifier, Stat De-id, based on support vector machines and local context (F-measure = 97% on PHI). Our representation of local context aids de-identification even when PHI include out-of-vocabulary words and even when PHI are ambiguous with non-PHI within the same corpus. Comparison of Stat De-id with a rule-based approach shows that Local context contributes more to de-identification than dictionaries combined with hand-tailored heuristics (F-measure = 85%). Comparison with two well-known named entity recognition (NER) systems, SNoW (F-measure = 94%) and IdentiFinder (F-measure = 36%), on five representative corpora show that when the language of documents is fragmented, a system with a relatively thorough representation of local context can be a more effective de-identifier than systems that combine (relatively simpler) local context with global context. Comparison with a Conditional Random Field De-identifier (CRFD), which utilizes global context in addition to the local context of Stat De-id, confirms this finding (F-measure=88%) and establishes that strengthening the representation of local context may be more beneficial for de-identification than complementing local with global context. (C) 2007 Elsevier B.V. All rights reserved.