Developing a scalable FHIR-based clinical data normalization pipeline for standardizing and integrating unstructured and structured electronic health record data

Developing a scalable FHIR-based clinical data normalization pipeline for standardizing and integrating unstructured and structured electronic health record data
复制标题

DOI:
10.1093/jamiaopen/ooz056
复制
发表时间:
2019-12-01
期刊:
影响因子:
2.1
通讯作者:
Jiang, Guoqian
Jiang, Guoqian
中科院分区:
其他
文献类型:
--
作者:
Hong, Na;Wen, Andrew;Jiang, Guoqian

文献摘要

被引文献

相似文献

目的:设计、开发和评估一个可扩展的临床数据规范化管道,利用HL 7快速医疗互操作资源(FHIR)规范对非结构化电子健康记录(EHR)数据进行标准化。方法:我们建立了一个基于FHIR的临床数据规范化管道NLP 2FHIR,主要包括:(1)一个基于FHIR类型系统的核心自然语言处理(NLP)引擎模块;(2)用于整合结构化数据的模块;以及(3)用于内容规范化的模块。我们使用马约诊所的非结构化EHR数据评估了FHIR建模能力,重点关注核心临床资源,如条件、程序、药物声明(包括药物)和家庭病史。我们构建了一个黄金标准重用注释语料库从以前的NLP projects.Results:共30个映射规则,62个规范化规则,和11个NLP特定的FHIR扩展被创建和实现在NLP 2FHIR管道。确定了需要整合来自每个临床资源的结构化数据的要素。对于各种FHIR元素表示,非结构化数据建模的性能达到了F分数,范围从0.69到0.99(病情为0.69-0.99;手术为0.75-0.84;用药说明为0.71-0.99;家族病史为0.75-0.95)。我们证明了NLP 2FHIR管道对于非结构化EHR数据建模和将结构化元素集成到模型中是可行的。这项工作的成果提供了基于标准的临床数据规范化工具,这对于实现便携式EHR驱动的表型分析和大规模数据分析是必不可少的,同时也为FHIR规范在处理非结构化临床数据方面的未来发展提供了有用的见解。
Objective: To design, develop, and evaluate a scalable clinical data normalization pipeline for standardizing unstructured electronic health record (EHR) data leveraging the HL7 Fast Healthcare Interoperability Resources (FHIR) specification.Methods: We established an FHIR-based clinical data normalization pipeline known as NLP2FHIR that mainly comprises: (1) a module for a core natural language processing (NLP) engine with an FHIR-based type system; (2) a module for integrating structured data; and (3) a module for content normalization. We evaluated the FHIR modeling capability focusing on core clinical resources such as Condition, Procedure, MedicationStatement (including Medication), and FamilyMemberHistory using Mayo Clinic's unstructured EHR data. We constructed a gold standard reusing annotation corpora from previous NLP projects.Results: A total of 30 mapping rules, 62 normalization rules, and 11 NLP-specific FHIR extensions were created and implemented in the NLP2FHIR pipeline. The elements that need to integrate structured data from each clinical resource were identified. The performance of unstructured data modeling achieved F scores ranging from 0.69 to 0.99 for various FHIR element representations (0.69-0.99 for Condition; 0.75-0.84 for Procedure; 0.71-0.99 for MedicationStatement; and 0.75-0.95 for FamilyMemberHistory).Conclusion: We demonstrated that the NLP2FHIR pipeline is feasible for modeling unstructured EHR data and integrating structured elements into the model. The outcomes of this work provide standards-based tools of clinical data normalization that is indispensable for enabling portable EHR-driven phenotyping and large-scale data analytics, as well as useful insights for future developments of the FHIR specifications with regard to handling unstructured clinical data.