Building bridges across electronic health record systems through inferred phenotypic topics

Building bridges across electronic health record systems through inferred phenotypic topics
复制标题

DOI:
10.1016/j.jbi.2015.03.011
复制
发表时间:
2015-06-01
影响因子:
4.5
通讯作者:
Malim, Bradley
Malim, Bradley
中科院分区:
医学3区
文献类型:
--
作者:
Chen, You;Ghosh, Joydeep;Malim, Bradley

文献摘要

被引文献

相似文献

目的:电子健康记录(EHR)中的数据越来越多地用于二次使用,从生物医学协会研究到比较有效性。为了进行大规模的研究,并以有意义的方式将知识从一个机构转移到另一个机构,我们需要协调这些系统中的表型。传统上,这已经通过经由标准化术语(例如账单代码)的表型的专家规范来实现。然而,这种方法可能会受到专家的经验和期望以及用于描述此类患者的词汇的影响。这项工作的目标是开发一个数据驱动的策略,(1)推断患者人群中的表型主题和(2)评估这些主题促进在不同的医疗保健systems.Methods的人群之间的映射的程度:我们适应一个生成主题建模策略,潜在的狄利克雷分配的基础上,推断表型主题。我们利用方差分析来评估从一个医疗保健系统到从另一个系统学到的主题的患者群体的投影。使用(1)主题的相似性,(2)患者群体在主题之间的稳定性,以及(3)主题在研究中心之间的可转移性来评估学习表型主题的一致性。我们使用来自两个地理上不同的医疗保健系统的四个月的住院患者数据评估了我们的方法:(1)西北纪念医院(NMH)和(2)范德比尔特大学医学中心(VUMC.Results):该方法从每个医疗保健系统学习了25个表型主题。两个网站匹配主题之间的平均余弦相似度为0.39,考虑到特征空间的高维性,这是一个非常高的值。VUMC和NMH患者的平均稳定性在两个站点的主题分别为0.988和0.812,如Pearson相关系数所测量的。此外,与标准临床术语相比,VUMC和NMH主题在表征两个研究中心的患者人群方面具有较小的方差(例如,ICD 9),这表明他们可能会更可靠地转移到医院systems.Conclusions:从EHR数据学习的表型主题可以更稳定和可转让的比计费代码的一般状态的患者人群的特征。这表明,基于EHR的研究可能能够在预测模型中汇集患者人群时利用这些表型主题作为变量。(C)2015 Elsevier Inc. All rights reserved.
Objective: Data in electronic health records (EHRs) is being increasingly leveraged for secondary uses, ranging from biomedical association studies to comparative effectiveness. To perform studies at scale and transfer knowledge from one institution to another in a meaningful way, we need to harmonize the phenotypes in such systems. Traditionally, this has been accomplished through expert specification of phenotypes via standardized terminologies, such as billing codes. However, this approach may be biased by the experience and expectations of the experts, as well as the vocabulary used to describe such patients. The goal of this work is to develop a data-driven strategy to (1) infer phenotypic topics within patient populations and (2) assess the degree to which such topics facilitate a mapping across populations in disparate healthcare systems.Methods: We adapt a generative topic modeling strategy, based on latent Dirichlet allocation, to infer phenotypic topics. We utilize a variance analysis to assess the projection of a patient population from one healthcare system onto the topics learned from another system. The consistency of learned phenotypic topics was evaluated using (1) the similarity of topics, (2) the stability of a patient population across topics, and (3) the transferability of a topic across sites. We evaluated our approaches using four months of inpatient data from two geographically distinct healthcare systems: (1) Northwestern Memorial Hospital (NMH) and (2) Vanderbilt University Medical Center (VUMC).Results: The method learned 25 phenotypic topics from each healthcare system. The average cosine similarity between matched topics across the two sites was 0.39, a remarkably high value given the very high dimensionality of the feature space. The average stability of VUMC and NMH patients across the topics of two sites was 0.988 and 0.812, respectively, as measured by the Pearson correlation coefficient. Also the VUMC and NMH topics have smaller variance of characterizing patient population of two sites than standard clinical terminologies (e.g., ICD9), suggesting they may be more reliably transferred across hospital systems.Conclusions: Phenotypic topics learned from EHR data can be more stable and transferable than billing codes for characterizing the general status of a patient population. This suggests that EHR-based research may be able to leverage such phenotypic topics as variables when pooling patient populations in predictive models. (C) 2015 Elsevier Inc. All rights reserved.