Development and evaluation of a de-identification procedure for a case register sourced from mental health electronic records.

Development and evaluation of a de-identification procedure for a case register sourced from mental health electronic records.
复制标题

DOI:
10.1186/1472-6947-13-71
复制
发表时间:
2013-07-11
影响因子:
3.5
通讯作者:
Callard F
Callard F
中科院分区:
医学3区
文献类型:
--
作者:
Fernandes AC;Cloete D;Broadbent MT;Hayes RD;Chang CK;Jackson RG;Roberts A;Tsang J;Soncul M;Liebscher J;Stewart R;Callard F

文献摘要

参考文献

被引文献

相似文献

电子健康记录(EHR)为健康研究提供了巨大的潜力,但也带来了数据治理的挑战。确保去识别是在未经事先同意的情况下使用EHR数据的先决条件。南伦敦和Maudsley NHS信托基金(SLaM)是欧洲最大的二级精神卫生保健提供者之一,它从EHR中开发了一个去识别的精神病病例登记册,临床记录交互式搜索(CRIS),用于二级研究。我们描述了一个定制的去识别算法用于创建注册的开发,实施和评估。它旨在使用输入到专用源字段中的患者标识符(PI)创建字典,然后在它们出现在医学文本中时识别,匹配和屏蔽它们(使用ZZZZZ)。我们认为这种方法将是有效的,因为PI在专用字段中的覆盖率很高,并且掩蔽的有效性与安全模型的元素相结合。我们进行了两个单独的性能测试i)以测试该算法在屏蔽在专用字段中输入的单个真实PI然后在文本中找到的性能(使用500个患者笔记)和ii)比较CRIS模式匹配算法与机器学习算法的性能,称为MITRE识别洗涤器工具包- MIST(使用70个患者笔记- 50个笔记用于训练,20个笔记用于测试)。我们还报告了潜在违规的任何发生率,定义为在同一患者的记录中(以及在50例患者的额外一组纵向记录中)发生3个或更多真实或明显的PI;我们考虑了尽管去识别,但推断信息的可能性。真实PI被掩蔽,精确度为98.8%,召回率为97.6%。正如预期的那样,由于电子健康记录中输入的拼写错误,潜在的PI确实出现了。我们发现了一个潜在的漏洞。在一个单独的性能测试中,使用不同的笔记集,CRIS产生了100%的准确率和88.5%的召回率,而MIST分别产生了95.1%和78.1%。我们将讨论如何克服现实的可能性-虽然概率低-通过实施的安全模型的潜在漏洞。CRIS是一个来自EHR的去身份化精神病学数据库,它保护患者的匿名性,并最大限度地利用可用于研究的数据。CRIS展示了将有效的去识别算法与精心设计的安全模型相结合的优势。本文提出了急需讨论的EHR去识别-特别是在有关的标准,以评估去识别,并考虑到背景下的去识别研究数据库时,评估违反保密患者信息的风险。
Electronic health records (EHRs) provide enormous potential for health research but also present data governance challenges. Ensuring de-identification is a pre-requisite for use of EHR data without prior consent. The South London and Maudsley NHS Trust (SLaM), one of the largest secondary mental healthcare providers in Europe, has developed, from its EHRs, a de-identified psychiatric case register, the Clinical Record Interactive Search (CRIS), for secondary research. We describe development, implementation and evaluation of a bespoke de-identification algorithm used to create the register. It is designed to create dictionaries using patient identifiers (PIs) entered into dedicated source fields and then identify, match and mask them (with ZZZZZ) when they appear in medical texts. We deemed this approach would be effective, given high coverage of PI in the dedicated fields and the effectiveness of the masking combined with elements of a security model. We conducted two separate performance tests i) to test performance of the algorithm in masking individual true PIs entered in dedicated fields and then found in text (using 500 patient notes) and ii) to compare the performance of the CRIS pattern matching algorithm with a machine learning algorithm, called the MITRE Identification Scrubber Toolkit – MIST (using 70 patient notes – 50 notes to train, 20 notes to test on). We also report any incidences of potential breaches, defined by occurrences of 3 or more true or apparent PIs in the same patient’s notes (and in an additional set of longitudinal notes for 50 patients); and we consider the possibility of inferring information despite de-identification. True PIs were masked with 98.8% precision and 97.6% recall. As anticipated, potential PIs did appear, owing to misspellings entered within the EHRs. We found one potential breach. In a separate performance test, with a different set of notes, CRIS yielded 100% precision and 88.5% recall, while MIST yielded a 95.1% and 78.1%, respectively. We discuss how we overcome the realistic possibility – albeit of low probability – of potential breaches through implementation of the security model. CRIS is a de-identified psychiatric database sourced from EHRs, which protects patient anonymity and maximises data available for research. CRIS demonstrates the advantage of combining an effective de-identification algorithm with a carefully designed security model. The paper advances much needed discussion of EHR de-identification – particularly in relation to criteria to assess de-identification, and considering the contexts of de-identified research databases when assessing the risk of breaches of confidential patient information.
DOI: 10.1186/1471-2288-10-70
发表时间: 2010-08-02
影响因子: 4
作者:
Meystre SM;Friedlin FJ;South BR;Shen S;Samore MH
通讯作者: Samore MH
DOI: 10.1186/1471-2105-12-s3-s2
发表时间: 2011-06-09
期刊: BMC bioinformatics
影响因子: 3
作者:
Benton A;Hill S;Ungar L;Chung A;Leonard C;Freeman C;Holmes JH
通讯作者: Holmes JH
DOI: 10.1186/1471-244x-9-51
发表时间: 2009-08-12
期刊: BMC PSYCHIATRY
影响因子: 4.4
作者:
Stewart, Robert;Soremekun, Mishael;Lovestone, Simon
通讯作者: Lovestone, Simon
DOI: 10.1186/1472-6947-8-32
发表时间: 2008-07-24
影响因子: 3.5
作者:
Neamatullah, Ishna;Douglass, Margaret M.;Clifford, Gari D.
通讯作者: Clifford, Gari D.
DOI: 10.1136/jamia.2010.004622
发表时间: 2011-01-01
影响因子: 6.4
作者:
Malin, Bradley;Benitez, Kathleen;Masys, Daniel
通讯作者: Masys, Daniel