Automated de-identification of free-text medical records

Automated de-identification of free-text medical records
复制标题

DOI:
10.1186/1472-6947-8-32
复制
发表时间:
2008-07-24
影响因子:
3.5
通讯作者:
Clifford, Gari D.
Clifford, Gari D.
中科院分区:
医学3区
文献类型:
--
作者:
Neamatullah, Ishna;Douglass, Margaret M.;Clifford, Gari D.

文献摘要

被引文献

相似文献

背景:基于文本的患者病历是医学研究中的重要资源。然而,为了保护病人的隐私,美国卫生部。S.健康保险携带和责任法案(HIPAA)要求在传播之前从医疗记录中删除受保护的健康信息(PHI)。手动去识别的大型医疗记录数据库是昂贵的,耗时的,容易出错,需要自动化的方法,大规模的,自动de-identification.Methods:我们描述了一个自动化的基于Perl的去识别软件包,通常是可用的大多数自由文本的医疗记录,例如。例如,在一个实施例中,该软件使用词汇查找表、正则表达式和简单的语法来定位HIPAA PHI和包括医生姓名和日期年份的扩展PHI集。为了开发去识别方法,我们组装了一个金标准语料库的重新识别护理笔记与真实的PHI取代现实的替代信息。该语料库包括2,434份护理记录,包含334,000个单词,以及从163份随机选择的患者记录中提取的总共1,779个PHI实例。这个黄金标准语料库被用来改进算法并测量其敏感性。为了在其开发中未使用的数据上测试该算法,我们构建了第二个测试语料库,其中包含1,836个护理笔记,包含296,400个单词。该算法的假阴性率进行了评估,使用此测试corpus.Results:性能评估的去识别软件的开发语料库产生的总体召回率为0.967,精度值为0.749,约0.002的辐射值。在测试语料库中,共发现90个假阴性实例,即每10万个单词中有27个,估计召回率为0.943。只有一个完整的日期和一个年龄超过89岁被错过了。没有病人的名字被遗漏,无论是corpus.Conclusion:我们已经开发了一个模式匹配去识别系统的基础上字典查找,正则表达式,和语法。对两组护理记录进行评价。S. hospital建议,在召回方面,该软件的性能优于单个人工去识别器(0.81),并且至少与两个人工去识别器的一致性(0.94)一样好。该系统目前正在调整,以去识别护理笔记和出院摘要中的PHI,但已足够通用,并可定制以处理任何格式的文本文件。虽然该算法的准确性很高,但它可能不足以用于公开传播医疗数据。因此,研究人员可以通过PhysioNet网站获得开源去识别软件和黄金标准的重新识别医疗记录语料库,以鼓励改进算法。
Background: Text-based patient medical records are a vital resource in medical research. In order to preserve patient confidentiality, however, the U. S. Health Insurance Portability and Accountability Act (HIPAA) requires that protected health information (PHI) be removed from medical records before they can be disseminated. Manual de-identification of large medical record databases is prohibitively expensive, time-consuming and prone to error, necessitating automatic methods for large-scale, automated de-identification.Methods: We describe an automated Perl-based de-identification software package that is generally usable on most free-text medical records, e. g., nursing notes, discharge summaries, X-ray reports, etc. The software uses lexical look-up tables, regular expressions, and simple heuristics to locate both HIPAA PHI, and an extended PHI set that includes doctors' names and years of dates. To develop the de-identification approach, we assembled a gold standard corpus of re-identified nursing notes with real PHI replaced by realistic surrogate information. This corpus consists of 2,434 nursing notes containing 334,000 words and a total of 1,779 instances of PHI taken from 163 randomly selected patient records. This gold standard corpus was used to refine the algorithm and measure its sensitivity. To test the algorithm on data not used in its development, we constructed a second test corpus of 1,836 nursing notes containing 296,400 words. The algorithm's false negative rate was evaluated using this test corpus.Results: Performance evaluation of the de-identification software on the development corpus yielded an overall recall of 0.967, precision value of 0.749, and fallout value of approximately 0.002. On the test corpus, a total of 90 instances of false negatives were found, or 27 per 100,000 word count, with an estimated recall of 0.943. Only one full date and one age over 89 were missed. No patient names were missed in either corpus.Conclusion: We have developed a pattern-matching de-identification system based on dictionary look-ups, regular expressions, and heuristics. Evaluation based on two different sets of nursing notes collected from a U. S. hospital suggests that, in terms of recall, the software out-performs a single human de-identifier (0.81) and performs at least as well as a consensus of two human de-identifiers (0.94). The system is currently tuned to de-identify PHI in nursing notes and discharge summaries but is sufficiently generalized and can be customized to handle text files of any format. Although the accuracy of the algorithm is high, it is probably insufficient to be used to publicly disseminate medical data. The open-source de-identification software and the gold standard re-identified corpus of medical records have therefore been made available to researchers via the PhysioNet website to encourage improvements in the algorithm.