Enabling qualitative research data sharing using a natural language processing pipeline for deidentification: moving beyond HIPAA Safe Harbor identifiers.

Enabling qualitative research data sharing using a natural language processing pipeline for deidentification: moving beyond HIPAA Safe Harbor identifiers.
复制标题

DOI:
10.1093/jamiaopen/ooab069
复制
发表时间:
2021-07
期刊:
影响因子:
2.1
通讯作者:
DuBois JM
DuBois JM
中科院分区:
其他
文献类型:
--
作者:
Gupta A;Lai A;Mozersky J;Ma X;Walsh H;DuBois JM

文献摘要

参考文献

被引文献

相似文献

共享卫生研究数据对于加速将研究转化为可影响卫生保健服务和成果的可操作知识至关重要。定性健康研究数据很少共享,因为文本去识别的挑战和参与者重新识别的潜在风险。在这里,我们建立并评估了一个框架,用于使用自动计算技术去除定性研究数据的标识,包括去除不被认为是HIPAA安全港(HSH)标识符,但可能在非结构化定性数据中发现的标识符。我们开发并验证了使用自动计算技术对定性研究数据进行去识别的管道。对不同类型的定性健康研究数据进行了深入的分析和定性回顾,以提供和评估自然语言处理(NLP)管道的发展,该管道使用命名实体识别、模式匹配、字典和正则表达式方法来去识别定性文本。我们收集了来自400多个定性研究数据文档的2个数据集,共120万字。我们创建了一个包含280K单词(70个文件)的金标准数据集来评估我们的去识别管道。定性数据中的大多数标识符是非hsh的,并且没有被现有系统捕获。我们的NLP去识别管道在两个数据集上的f1得分一致,为0.90。本研究的结果表明,NLP方法可以用于识别HSH标识符和非HSH标识符。鉴于新的国家卫生研究院(NIH)数据共享授权,帮助研究人员对定性数据进行去识别的自动化工具将变得越来越重要。
Sharing health research data is essential for accelerating the translation of research into actionable knowledge that can impact health care services and outcomes. Qualitative health research data are rarely shared due to the challenge of deidentifying text and the potential risks of participant reidentification. Here, we establish and evaluate a framework for deidentifying qualitative research data using automated computational techniques including removal of identifiers that are not considered HIPAA Safe Harbor (HSH) identifiers but are likely to be found in unstructured qualitative data. We developed and validated a pipeline for deidentifying qualitative research data using automated computational techniques. An in-depth analysis and qualitative review of different types of qualitative health research data were conducted to inform and evaluate the development of a natural language processing (NLP) pipeline using named-entity recognition, pattern matching, dictionary, and regular expression methods to deidentify qualitative texts. We collected 2 datasets with 1.2 million words derived from over 400 qualitative research data documents. We created a gold-standard dataset with 280K words (70 files) to evaluate our deidentification pipeline. The majority of identifiers in qualitative data are non-HSH and not captured by existing systems. Our NLP deidentification pipeline had a consistent F1-score of ∼0.90 for both datasets. The results of this study demonstrate that NLP methods can be used to identify both HSH identifiers and non-HSH identifiers. Automated tools to assist researchers with the deidentification of qualitative data will be increasingly important given the new National Institutes of Health (NIH) data-sharing mandate.
DOI: 10.1177/1468794114550439
发表时间: 2015-10-01
影响因子: 3.6
作者:
Saunders, Benjamin;Kitzinger, Jenny;Kitzinger, Celia
通讯作者: Kitzinger, Celia
DOI: 10.1186/1472-6947-8-32
发表时间: 2008-07-24
影响因子: 3.5
作者:
Neamatullah, Ishna;Douglass, Margaret M.;Clifford, Gari D.
通讯作者: Clifford, Gari D.
DOI: 10.29173/iq952
发表时间: 2020-01-08
期刊: IASSIST quarterly
影响因子: --
作者:
Mozersky, Jessica;Walsh, Heidi;DuBois, James M
通讯作者: DuBois, James M
通过循环神经网络和条件随机场对临床记录进行去识别
DOI: 10.1016/j.jbi.2017.05.023
发表时间: 2017-11
影响因子: 4.5
作者:
Liu Z;Tang B;Wang X;Chen Q
通讯作者: Chen Q
DOI: 10.1038/s41746-020-0258-y
发表时间: 2020-04-14
影响因子: 15.2
作者:
Norgeot, Beau;Muenzen, Kathleen;Butte, Atul J.
通讯作者: Butte, Atul J.