A corpus-driven standardization framework for encoding clinical problems with HL7 FHIR.

A corpus-driven standardization framework for encoding clinical problems with HL7 FHIR.
复制标题

DOI:
10.1016/j.jbi.2020.103541
复制
发表时间:
2020-10
影响因子:
4.5
通讯作者:
Liu H
Liu H
中科院分区:
医学3区
文献类型:
--
作者:
Peterson KJ;Jiang G;Liu H

文献摘要

参考文献

被引文献

相似文献

自由文本问题描述是对患者诊断和问题的简要说明,通常在问题列表和医疗记录的其他重要区域中找到。这些紧凑的表示通常表示复杂和细微差别的医疗条件,这使得它们的语义具有完全捕获和标准化的挑战性。在这项研究中,我们描述了一个框架,用于将自由文本问题描述转换为标准化的Health Level 7(HL7)快速医疗互操作性资源(FHIR)模型。该方法利用特定于领域的依赖关系解析器、来自Transformers(BERT)自然语言模型的双向编码器表示和CUI2vec统一医疗语言系统(UMLS)概念向量的组合,将从自由文本问题描述中提取的概念与结构化FHIR模型对齐。使用神经网络分类模型对概念之间的13种关系类型进行分类,便于映射到冷杉条件资源。我们使用数据编程,这是一种弱监督方法,以消除对人工标注训练语料库的需要。Shapley值是一种量化贡献的机制,用于解释模型特征的影响。我们发现,我们的方法确定了焦点概念,或问题描述的主要临床问题,F1得分为0.95。从焦点到其他修饰概念的关系被提取出来,F1得分为0.90。在对关系进行分类时,我们的模型达到了0.89的加权平均F1分数,使得属性能够准确地映射到HL7 FHIR模型中。我们还发现,正如Shapley值分析所显示的那样,BERT输入表示主要对分类器决策做出贡献。
Free-text problem descriptions are brief explanations of patient diagnoses and issues, commonly found in problem lists and other prominent areas of the medical record. These compact representations often express complex and nuanced medical conditions, making their semantics challenging to fully capture and standardize. In this study, we describe a framework for transforming free-text problem descriptions into standardized Health Level 7 (HL7) Fast Healthcare Interoperability Resources (FHIR) models. This approach leverages a combination of domain-specific dependency parsers, Bidirectional Encoder Representations from Transformers (BERT) natural language models, and cui2vec Unified Medical Language System (UMLS) concept vectors to align extracted concepts from free-text problem descriptions into structured FHIR models. A neural network classification model is used to classify thirteen relationship types between concepts, facilitating mapping to the FHIR Condition resource. We use data programming, a weak supervision approach, to eliminate the need for a manually annotated training corpus. Shapley values, a mechanism to quantify contribution, are used to interpret the impact of model features. We found that our methods identified the focus concept, or primary clinical concern of the problem description, with an F1 score of 0.95. Relationships from the focus to other modifying concepts were extracted with an F1 score of 0.90. When classifying relationships, our model achieved a 0.89 weighted average F1 score, enabling accurate mapping of attributes into HL7 FHIR models. We also found that the BERT input representation predominantly contributed to the classifier decision as shown by the Shapley values analysis.
DOI: 10.1370/afm.2121
发表时间: 2017-09-01
影响因子: 4.4
作者:
Arndt, Brian G.;Beasley, John W.;Gilchrist, Valerie J.
通讯作者: Gilchrist, Valerie J.
DOI: 10.1093/jamiaopen/ooz056
发表时间: 2019-12-01
期刊: JAMIA OPEN
影响因子: 2.1
作者:
Hong, Na;Wen, Andrew;Jiang, Guoqian
通讯作者: Jiang, Guoqian
DOI: 10.1093/bioinformatics/btl616
发表时间: 2007-02-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Fundel, Katrin;Kueffner, Robert;Zimmer, Ralf
通讯作者: Zimmer, Ralf
DOI: 10.1080/19312450709336664
发表时间: 2007-01-01
影响因子: 11.4
作者:
Hayes, Andrew F.;Krippendorff, Klaus
通讯作者: Krippendorff, Klaus
DOI: 10.1093/nar/gkh061
发表时间: 2004-01-01
影响因子: 14.9
作者:
Bodenreider, O
通讯作者: Bodenreider, O