Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data.

Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data.
复制标题

DOI:
10.1038/s41746-021-00519-z
复制
发表时间:
2021-10-27
影响因子:
15.2
通讯作者:
VA Million Veteran Program
VA Million Veteran Program
中科院分区:
医学1区
文献类型:
--
作者:
Hong C;Rush E;Liu M;Zhou D;Sun J;Sonabend A;Castro VM;Schubert P;Panickan VA;Cai T;Costa L;He Z;Link N;Hauser R;Gaziano JM;Murphy SN;Ostrouchov G;Ho YL;Begoli E;Lu J;Cho K;Liao KP;Cai T;VA Million Veteran Program

文献摘要

参考文献

被引文献

相似文献

电子健康记录(EHR)系统的日益普及为转化研究创造了巨大的潜力。然而,由于可用的编码数量众多,很难知道与表型相关的所有相关编码。传统的数据挖掘方法通常需要使用患者级别的数据,这阻碍了跨机构共享数据的能力。在这个项目中,我们证明了多中心大规模代码嵌入可以用来有效地识别与感兴趣的疾病相关的相关特征。我们为来自两个大型医疗中心的电子病历的广泛的编码概念构建了大规模的代码嵌入。我们开发了基于稀疏嵌入回归(KESER)的知识提取方法,用于特征选择和综合网络分析。我们评估了代码嵌入的质量,并评估了KESER在8种疾病的特征选择中的性能。此外,我们结合两个机构的嵌入数据,开发了一个集成的临床知识图谱。KESER选择的特征与领域专家生成的编码数据列表相比是全面的。通过KESER识别的特性与基于手动选择的特性或基于患者级数据构建的特性的性能相当。与使用单一机构数据确定的知识图相比,使用综合分析创建的知识图更准确地确定了疾病-疾病和疾病-药物对。利用KESER对编码嵌入进行分析,可以有效地揭示临床知识并推断编码概念之间的相关性。KESER在个人分析中绕过了对患者水平数据的需求,为使用电子病历数据进行多中心研究提供了重大进展。
The increasing availability of electronic health record (EHR) systems has created enormous potential for translational research. However, it is difficult to know all the relevant codes related to a phenotype due to the large number of codes available. Traditional data mining approaches often require the use of patient-level data, which hinders the ability to share data across institutions. In this project, we demonstrate that multi-center large-scale code embeddings can be used to efficiently identify relevant features related to a disease of interest. We constructed large-scale code embeddings for a wide range of codified concepts from EHRs from two large medical centers. We developed knowledge extraction via sparse embedding regression (KESER) for feature selection and integrative network analysis. We evaluated the quality of the code embeddings and assessed the performance of KESER in feature selection for eight diseases. Besides, we developed an integrated clinical knowledge map combining embedding data from both institutions. The features selected by KESER were comprehensive compared to lists of codified data generated by domain experts. Features identified via KESER resulted in comparable performance to those built upon features selected manually or with patient-level data. The knowledge map created using an integrative analysis identified disease-disease and disease-drug pairs more accurately compared to those identified using single institution data. Analysis of code embeddings via KESER can effectively reveal clinical knowledge and infer relatedness among codified concepts. KESER bypasses the need for patient-level data in individual analyses providing a significant advance in enabling multi-center studies using EHR data.
DOI: 10.1038/nbt.2749
发表时间: 2013-12
影响因子: 46.9
作者:
通讯作者: --
DOI: 10.3390/jpm6010002
发表时间: 2016-03-01
影响因子: --
作者:
Karlson, Elizabeth W.;Boutin, Natalie T.;Allen, Nicole L.
通讯作者: Allen, Nicole L.
DOI: 10.1093/jamia/ocw112
发表时间: 2017-03-01
期刊: Journal of the American Medical Informatics Association : JAMIA
影响因子: --
作者:
Choi E;Schuetz A;Stewart WF;Sun J
通讯作者: Sun J
DOI: 10.1373/49.4.624
发表时间: 2003-04-01
期刊: CLINICAL CHEMISTRY
影响因子: 9.3
作者:
McDonald, CJ;Huff, SM;Maloney, P
通讯作者: Maloney, P
DOI: 10.1007/s00392-016-1025-6
发表时间: 2017-01
期刊: Clinical research in cardiology : official journal of the German Cardiac Society
影响因子: --
作者:
Cowie MR;Blomster JI;Curtis LH;Duclaux S;Ford I;Fritz F;Goldman S;Janmohamed S;Kreuzer J;Leenay M;Michel A;Ong S;Pell JP;Southworth MR;Stough WG;Thoenes M;Zannad F;Zalewski A
通讯作者: Zalewski A