Strategies for De-identification and Anonymization of Electronic Health Record Data for Use in Multicenter Research Studies

Strategies for De-identification and Anonymization of Electronic Health Record Data for Use in Multicenter Research Studies
复制标题

DOI:
10.1097/mlr.0b013e3182585355
复制
发表时间:
2012-07-01
期刊:
影响因子:
3
通讯作者:
Griffin, Kara
Griffin, Kara
中科院分区:
医学3区
文献类型:
--
作者:
Kushida, Clete A.;Nichols, Deborah A.;Griffin, Kara

文献摘要

被引文献

相似文献

背景:去标识化和匿名化是用于在电子健康记录数据中删除患者标识符的策略。考虑到需要在多个环境和机构之间共享电子健康记录数据,同时保护患者隐私,在多中心研究中使用这些策略至关重要。方法:采用“去身份”、“去身份”、“去身份”、“去身份”、“匿名”、“匿名化”、“数据清洗”、“文本清洗”等关键词进行系统文献检索。检索截止到2011年6月30日,涉及6个不同的常用文献数据库。共鉴定出1798条预期引用,94篇全文文章符合评审标准,获得相应的文章。检索结果补充了26篇额外的全文文章;共审查了120篇全文文章。结果:45篇文章的最终样本符合纳入标准进行审查和讨论。文章分为文本、图像和生物样本三类。对于基于文本的策略,方法被分为启发式、词汇和基于模式的系统与基于统计学习的系统。对于图像,描述了去识别摄影面部图像和磁共振图像数据的方法。对于生物样本,讨论了管理与这些样本相关的标识符的方法,特别是关于满足机构审查委员会根据共同规则豁免所需的匿名化要求。结论:当前的去识别策略有其局限性,基于统计学习的系统在自由文本去识别方面比其他方法具有明显的优势。真正的匿名化是具有挑战性的,在数据集的去识别和遗传信息的保护方面需要进一步的工作。
Background: De-identification and anonymization are strategies that are used to remove patient identifiers in electronic health record data. The use of these strategies in multicenter research studies is paramount in importance, given the need to share electronic health record data across multiple environments and institutions while safeguarding patient privacy.Methods: Systematic literature search using keywords of de-identify, deidentify, de-identification, deidentification, anonymize, anonymization, data scrubbing, and text scrubbing. Search was conducted up to June 30, 2011 and involved 6 different common literature databases. A total of 1798 prospective citations were identified, and 94 full-text articles met the criteria for review and the corresponding articles were obtained. Search results were supplemented by review of 26 additional full-text articles; a total of 120 full-text articles were reviewed.Results: A final sample of 45 articles met inclusion criteria for review and discussion. Articles were grouped into text, images, and biological sample categories. For text-based strategies, the approaches were segregated into heuristic, lexical, and pattern-based systems versus statistical learning-based systems. For images, approaches that de-identified photographic facial images and magnetic resonance image data were described. For biological samples, approaches that managed the identifiers linked with these samples were discussed, particularly with respect to meeting the anonymization requirements needed for Institutional Review Board exemption under the Common Rule.Conclusions: Current de-identification strategies have their limitations, and statistical learning-based systems have distinct advantages over other approaches for the de-identification of free text. True anonymization is challenging, and further work is needed in the areas of de-identification of datasets and protection of genetic information.