De-identification of patient notes with recurrent neural networks

De-identification of patient notes with recurrent neural networks
复制标题

DOI:
10.1093/jamia/ocw156
复制
发表时间:
2017-05-01
影响因子:
6.4
通讯作者:
Szolovits, Peter
Szolovits, Peter
中科院分区:
管理学2区
文献类型:
--
作者:
Dernoncourt, Franck;Lee, Ji Young;Szolovits, Peter

文献摘要

被引文献

相似文献

目的:电子健康记录(EHR)中的患者记录可能包含医学调查的关键信息。然而,绝大多数医学调查人员只能访问去身份化的笔记,以保护患者的机密性。在美国,《健康保险流通与责任法案》(Health Insurance Portability and Accountability Act,HIPAA)定义了18种受保护的健康信息,这些信息需要被删除以去识别患者记录。考虑到电子健康记录数据库的规模、能够访问非去识别笔记的研究人员数量有限以及人类注释者的频繁错误,手动去识别是不切实际的。一个可靠的自动去识别系统将因此是高value.Materials和方法:我们介绍了第一个去识别系统的基础上人工神经网络(ANN),它不需要手工制作的功能或规则,不像现有的系统。我们在两个数据集上比较了该系统与最先进系统的性能:i2b2 2014去识别挑战数据集,这是最大的公开去识别数据集,以及MIMIC去识别数据集,我们组装的数据集是i2b2 2014数据集的两倍。结果:我们的ANN模型优于最先进的系统。它在i2b2 2014数据集上的F1得分为97.85,召回率为97.38,精度为98.32,在MIMIC去识别数据集上的F1得分为99.23,召回率为99.25,精度为99.21。我们的研究结果支持使用人工神经网络去识别病人的笔记,因为他们表现出更好的性能比以前公布的系统,同时不需要手动功能工程。
Objective: Patient notes in electronic health records (EHRs) may contain critical information for medical investigations. However, the vast majority of medical investigators can only access de-identified notes, in order to protect the confidentiality of patients. In the United States, the Health Insurance Portability and Accountability Act (HIPAA) defines 18 types of protected health information that needs to be removed to de-identify patient notes. Manual de-identification is impractical given the size of electronic health record databases, the limited number of researchers with access to non-de-identified notes, and the frequent mistakes of human annotators. A reliable automated de-identification system would consequently be of high value.Materials and Methods: We introduce the first de-identification system based on artificial neural networks (ANNs), which requires no handcrafted features or rules, unlike existing systems. We compare the performance of the system with state-of-the-art systems on two datasets: the i2b2 2014 de-identification challenge dataset, which is the largest publicly available de-identification dataset, and the MIMIC de-identification dataset, which we assembled and is twice as large as the i2b2 2014 dataset.Results: Our ANN model outperforms the state-of-the-art systems. It yields an F1-score of 97.85 on the i2b2 2014 dataset, with a recall of 97.38 and a precision of 98.32, and an F1-score of 99.23 on the MIMIC de-identification dataset, with a recall of 99.25 and a precision of 99.21.Conclusion: Our findings support the use of ANNs for de-identification of patient notes, as they show better performance than previously published systems while requiring no manual feature engineering.