Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonization

Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonization
复制标题

DOI:
10.1016/j.jbi.2022.104147
复制
发表时间:
2022-07
影响因子:
4.5
通讯作者:
D. Zhou;Ziming Gan;Xu Shi;Alina Patwari;E. Rush;Clara-Lea Bonzel;V. A. Panickan;C. Hong;Y. Ho;T. Cai;L. Costa;Xiaoou Li;V. Castro;S. Murphy;G. Brat;G. Weber;P. Avillach;J. Gaziano;Kelly Cho;K. Liao;Junwei Lu;Tianxi Cai
D. Zhou;Ziming Gan;Xu Shi;Alina Patwari;E. Rush;Clara-Lea Bonzel;V. A. Panickan;C. Hong;Y. Ho;T. Cai;L. Costa;Xiaoou Li;V. Castro;S. Murphy;G. Brat;G. Weber;P. Avillach;J. Gaziano;Kelly Cho;K. Liao;Junwei Lu;Tianxi Cai
中科院分区:
医学3区
文献类型:
--
作者:
D. Zhou;Ziming Gan;Xu Shi;Alina Patwari;E. Rush;Clara-Lea Bonzel;V. A. Panickan;C. Hong;Y. Ho;T. Cai;L. Costa;Xiaoou Li;V. Castro;S. Murphy;G. Brat;G. Weber;P. Avillach;J. Gaziano;Kelly Cho;K. Liao;Junwei Lu;Tianxi Cai

文献摘要

相似文献

目的电子健康记录(EHR)数据的日益普及为多机构EHR的综合分析提供了机会,从而产生可概括的知识。这种综合分析的一个关键障碍是,由于编码差异,不同机构之间缺乏语义互操作性。提出了一种多视图不完全知识图集成(MIKGI)算法,将来自多个来源的信息与部分重叠的EHR概念代码进行集成,从而实现医疗系统之间的翻译。方法MIKGI算法结合了来自(I)从每个EHR系统中的医疗代码的共现模式训练的嵌入和(Ii)由自对齐预训练BERT(SAPBERT)算法获得的所有医疗代码的文本串的语义嵌入的知识图信息。由于医疗保健系统中编码的异质性,每个EHR来源都提供了可用代码的部分覆盖。MIKGI通过最小化球面损失函数来合成从这些多源嵌入得到的不完全知识图,该球面损失函数结合了从所有可用源计算的嵌入的成对方向相似性。MIKGI为所有EHR代码输出协调的语义嵌入向量,这提高了嵌入的质量,并允许直接评估来自多个医疗系统的任何一对代码之间的相似性和关联性。结果使用来自退伍军人事务部(VA)医疗保健和麻省总医院(MGB)的EHR同现数据,MIKGI算法为各种下游任务生成高质量的嵌入,包括检测已知的相似或相关实体对,以及将VA本地代码映射到MGB使用的相关EHR代码。基于MIKGI训练的嵌入体的余弦相似性,AUC值分别为0.918和0.809。对于跨机构医疗代码映射,将VA的药物代码映射到MGB的RxNorm药物代码时,前1名和前5名的准确率分别为91.0%和97.5%;将VA的本地实验室代码映射到LOINC层次时,前1名和前5名的准确率分别为59.1%和75.8%。当训练500个标签时,实验室编码映射的准确率分别达到了77.7%和87.9%。MIKGI在为所需的实验室测试选择VA本地实验室代码以及为COVID EHR研究选择新冠肺炎相关功能方面也取得了最佳表现。结论MIKGI算法可以有效地整合生物医学文本和电子病历数据中的不完整摘要数据,为知识图谱建模和电子病历代码的跨机构翻译生成协调的嵌入。
ObjectiveThe growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose aMultiviewIncompleteKnowledgeGraphIntegration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems.MethodsThe MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems.ResultsWith EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks.ConclusionsThe proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes.