NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
批准号:
10268072
负责人:
Mehmet Kayaalp
金额:
$84.51万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AddressAlgorithmsAreaArtificial IntelligenceClinicalClinical DataClinical ResearchComputational LinguisticsComputer softwareDatabasesDictionaryFloridaFutureGoalsGuidelinesHealthInternationalKnowledgeLabelLaboratoriesLawsMethodsNCI Center for Cancer ResearchNamesNational Cancer InstituteOutputPathology ReportPatientsPattern RecognitionPerformancePersonally Identifiable InformationPoliciesPrivacyProcessProviderPythonsRegulationReportingRiskSEER ProgramScientistShapesSocial Security NumberSoftware DesignSystemTechniquesTechnologyTelephoneTestingTextUnited States National Institutes of HealthUnited States National Library of MedicineUniversitiesVoicebasecohortcollegedata de-identificationflexibilityinterestmedical schoolsneoplasm registrypatient privacypreservationrepositorytool
中文摘要
叙述性临床报告包含一套丰富的临床知识,对临床研究可能是无价的。然而,它们也可能包含个人身份信息(PII),使这些临床报告归类为个人身份信息,这与使用限制和隐私风险有关。计算去识别试图删除此类叙事文本中的所有PII实例,以生成去识别文档,这些文档将不再被归类为PHI,并且可以在较少约束和几乎没有隐私风险的情况下用于临床研究。计算去识别使用人工智能方法,包括模式识别和计算语言技术来识别文本中表示PII的单词和其他字母数字令牌(例如,姓名、地址、电话和社会安全号码),并将其替换为诸如NAME和ADDRESS之类的标签。这样既保护了患者隐私,又保留了临床知识。
英文摘要
Narrative clinical reports contain a rich set of clinical knowledge that could be invaluable for clinical research. However, they may also contain personally identifiable information (PII) that make those clinical reports classified as PHI, which is associated with use restrictions and risks to privacy. Computational de-identification seeks to remove all instances of PII in such narrative text in order to produce de-identified documents, which would no longer be classified as PHI and can be used in clinical research with fewer constraints and with almost no risk to privacy. Computational de-identification uses artificial intelligence methods including pattern recognition and computational linguistic techniques to recognize words and other alphanumeric tokens denoting PII (e.g., names, addresses, and telephone and social security numbers) in the text, and replace them with labels such as NAME and ADDRESS. In this way, both patient privacy is protected and clinical knowledge is preserved.
After exploring existing de-identification tools, the U.S. National Library of Medicine (NLM) began developing a new software application called NLM Scrubber, which is capable of de-identifying many types of clinical reports with high accuracy. The software design is based on both deterministic and probabilistic artificial intelligence methods utilizing large dictionaries of personal names, addresses, and organizations. The application accepts narrative reports in plain text or in HL7 format. When the input reports are formatted as HL7 messages, the application software leverages patient information embedded in HL7 segments to find such information in the text portion of the HL7 message.
NLM Scrubber has been downloaded by a number of organizations for testing and use, including IBM, Google, Fred Hutch Cancer Research Center, Harvard Medical School, Florida International University, University of Bristol, University College Dublin, and Oak Ridge National Laboratory. National Cancer Institute (NCI) along with the state cancer registries working with NCI have also been voiced their interests in using NLM Scrubber to de-identify narrative pathology reports in Surveillance, Epidemiology and End Results (SEER) database.
Our current focus is making NLM-Scrubber answer various needs of clinical scientists and clinical data managers without a deep understanding of the underlying technology. In this term, we started moving our codebase from Perl to a C++/Python-based system, which will provide us more flexibility, higher performance and scalability for future functionalities.
While NLM-Scrubber can be used for de-identifying all clinical reports repository-wide, it can also be used in various modes tailored to the user and their context, including on-demand cohort-specific de-identification and de-identification with patient and provider identifiers. NLM-Scrubber enables clinical scientists, the users of de-identified data, to be part of the process and shape the de-identification output based on their needs.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Challenges and Insights in Using HIPAA Privacy Rule for Clinical Text Annotation.
使用 HIPAA 隐私规则进行临床文本注释的挑战和见解。
DOI:
--
发表时间:
2015
期刊:
AMIA ... Annual Symposium proceedings. AMIA Symposium
影响因子:
--
作者:
[Kayaalp,Mehmet, Browne,AllenC, Sagan,Pamela, McGee,Tyne, McDonald,ClementJ]
通讯作者:
McDonald,ClementJ
DOI:
10.4274/balkanmedj.2017.0966
发表时间:
2018-01-20
期刊:
Balkan medical journal
影响因子:
3
作者:
[Kayaalp M]
通讯作者:
Kayaalp M
Piloting a deceased subject integrated data repository and protecting privacy of relatives.
试点已故受试者综合数据存储库并保护亲属隐私。
DOI:
--
发表时间:
2014
期刊:
AMIA ... Annual Symposium proceedings / AMIA Symposium. AMIA Symposium
影响因子:
--
作者:
[Huser,Vojtech, Kayaalp,Mehmet, Dodd,ZeynoA, Cimino,JamesJ]
通讯作者:
Cimino,JamesJ
De-identification of Address, Date, and Alphanumeric Identifiers in Narrative Clinical Reports
叙述性临床报告中地址、日期和字母数字标识符的去识别化
DOI:
--
发表时间:
2014
期刊:
American Medical Informatics Association Annual Symposium
影响因子:
--
作者:
[M. Kayaalp, Allen C. Browne, Zeyno A. Dodd, P. Sagan, C. McDonald]
通讯作者:
C. McDonald
DOI:
10.4103/2153-3539.117450
发表时间:
2013
期刊:
Journal of pathology informatics
影响因子:
--
作者:
[Kang YS, Kayaalp M]
通讯作者:
Kayaalp M
共 6 条
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8558114
-
项目类别:
-
资助金额:$34.93万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
-
批准号:9554455
-
项目类别:
-
资助金额:$46.34万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8344957
-
项目类别:
-
资助金额:$33.32万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8158053
-
项目类别:
-
资助金额:$31.71万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8943232
-
项目类别:
-
资助金额:$38.53万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
海外基金