Comparative effectiveness of medical concept embedding for feature engineering in phenotyping.

Comparative effectiveness of medical concept embedding for feature engineering in phenotyping.
复制标题

DOI:
10.1093/jamiaopen/ooab028
复制
发表时间:
2021-04
期刊:
影响因子:
2.1
通讯作者:
Weng C
Weng C
中科院分区:
其他
文献类型:
--
作者:
Lee J;Liu C;Kim JH;Butler A;Shang N;Pang C;Natarajan K;Ryan P;Ta C;Weng C

文献摘要

参考文献

被引文献

相似文献

特征工程是表型分析的主要瓶颈。正确学习的医学概念嵌入(MCE)可以捕获医学概念的语义,因此对于检索表型任务中的相关医学特征非常有用。我们比较了从知识图和电子医疗记录 (EHR) 数据中学习的 MCE 在检索表型任务的相关医学特征方面的有效性。我们使用 2 个数据源实现了 5 种嵌入方法,包括 node2vec、奇异值分解(SVD)、LINE、skip-gram 和 GloVe:(1)从观察性医疗结果合作伙伴关系(OMOP)通用数据模型获得的知识图; (2) 从哥伦比亚大学欧文医学中心 (CUIMC) 的 OMOP 兼容电子健康记录 (EHR) 获取的患者级别数据。我们使用电子病历和基因组学 (eMERGE) 网络开发和验证的表型及其相关概念来评估学习的 MCE 在检索表型相关概念方面的表现。基于单个和多个种子概念检索表型相关概念的 Hits@k% 用于评估 MCE。在所有MCE中,使用node2vec和知识图学习的MCE表现出最好的性能。在基于知识图和 EHR 数据的 MCE 中,使用带有知识图的 Node2vec 学习的 MCE 和使用带有 EHR 数据的 GloVe 学习的 MCE 分别优于其他 MCE。 MCE 支持可扩展的特征工程任务,从而促进表型分析。根据当前的表型分析实践,通过使用由医学概念之间的层次关系构建的知识图来学习的 MCE 优于通过使用 EHR 数据来学习的 MCE。
Feature engineering is a major bottleneck in phenotyping. Properly learned medical concept embeddings (MCEs) capture the semantics of medical concepts, thus are useful for retrieving relevant medical features in phenotyping tasks. We compared the effectiveness of MCEs learned from knowledge graphs and electronic healthcare records (EHR) data in retrieving relevant medical features for phenotyping tasks. We implemented 5 embedding methods including node2vec, singular value decomposition (SVD), LINE, skip-gram, and GloVe with 2 data sources: (1) knowledge graphs obtained from the observational medical outcomes partnership (OMOP) common data model; and (2) patient-level data obtained from the OMOP compatible electronic health records (EHR) from Columbia University Irving Medical Center (CUIMC). We used phenotypes with their relevant concepts developed and validated by the electronic medical records and genomics (eMERGE) network to evaluate the performance of learned MCEs in retrieving phenotype-relevant concepts. Hits@k% in retrieving phenotype-relevant concepts based on a single and multiple seed concept(s) was used to evaluate MCEs. Among all MCEs, MCEs learned by using node2vec with knowledge graphs showed the best performance. Of MCEs based on knowledge graphs and EHR data, MCEs learned by using node2vec with knowledge graphs and MCEs learned by using GloVe with EHR data outperforms other MCEs, respectively. MCE enables scalable feature engineering tasks, thereby facilitating phenotyping. Based on current phenotyping practices, MCEs learned by using knowledge graphs constructed by hierarchical relationships among medical concepts outperformed MCEs learned by using EHR data.
DOI: 10.1111/biom.12987
发表时间: 2019-03-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
Gronsbell, Jessica;Minnier, Jessica;Cai, Tianxi
通讯作者: Cai, Tianxi
DOI: 10.1016/j.jbi.2019.103293
发表时间: 2019-11-01
影响因子: 4.5
作者:
Shang, Ning;Liu, Cong;Weng, Chunhua
通讯作者: Weng, Chunhua
DOI: 10.1145/2939672.2939754
发表时间: 2016-08
期刊: KDD : proceedings. International Conference on Knowledge Discovery & Data Mining
影响因子: --
作者:
Grover A;Leskovec J
通讯作者: Leskovec J
DOI: 10.1016/j.jbi.2019.103246
发表时间: 2019-08-01
影响因子: 4.5
作者:
Shen, Feichen;Peng, Suyuan;Liu, Hongfang
通讯作者: Liu, Hongfang
DOI: 10.1146/annurev-biodatasci-080917-013315
发表时间: 2018-01-01
期刊: ANNUAL REVIEW OF BIOMEDICAL DATA SCIENCE, VOL 1
影响因子: --
作者:
Banda, Juan M.;Seneviratne, Martin;Shah, Nigam H.
通讯作者: Shah, Nigam H.