HPO2Vec+: Leveraging heterogeneous knowledge resources to enrich node embeddings for the Human Phenotype Ontology

HPO2Vec+: Leveraging heterogeneous knowledge resources to enrich node embeddings for the Human Phenotype Ontology
复制标题

DOI:
10.1016/j.jbi.2019.103246
复制
发表时间:
2019-08-01
影响因子:
4.5
通讯作者:
Liu, Hongfang
Liu, Hongfang
中科院分区:
医学3区
文献类型:
--
作者:
Shen, Feichen;Peng, Suyuan;Liu, Hongfang

文献摘要

被引文献

相似文献

背景:在精准医学中,深度表型分析被定义为对表型异常进行精确、全面的分析,旨在更好地了解疾病的自然史及其基因型-表型关联。将精准医学转化为临床实践时,检测表型相关性是一项重要任务,尤其是基于深度表型分析的患者分层任务。在我们之前的工作中,我们为人类表型本体(HPO)开发了节点嵌入,以协助结合分布式语义表示的表型相关性测量。然而,派生的 HPO 嵌入仅保留节点之间 IS-A 关系的分布式表示,从而妨碍了充分探索图的能力。 方法:在本研究中,我们开发了一个框架 HPO2Vec +,用异构知识资源(即 DECIPHER、OMIM 和 Orphanet)丰富生成的 HPO 嵌入,以检测表型相关性。具体来说,我们解析了这三个资源中包含的疾病表型关联,以丰富 HPO 中表型节点之间的非遗传关系。为了生成 HPO 的节点嵌入,应用 no-de2vec 基于随机游走对丰富的 HPO 图执行节点采样,然后对采样的节点进行特征学习以生成丰富的节点嵌入。基于不同的图结构生成了四种 HPO 嵌入,我们在下文中将其标记为 HPOEmb-Original、HPOEmb-DECIPHER、HPOEmb-OMIM 和 HPOEmb-Orphanet。我们通过具有四个边缘嵌入操作和六种机器学习算法的 HPO 链接预测任务定量评估了派生的嵌入。然后使用 Mayo Clinic 收集的电子健康记录 (EHR) 评估最终的最佳嵌入,以对 10 种罕见疾病进行患者分层。我们通过可视化表型簇并对罕见疾病原发性高草酸尿症 (PH) 进行用例研究来定性评估我们的框架,任务是根据 22 个带注释的 PH 相关表型推断相关表型。结果:定量链接预测任务显示 HPOEmb-Orphanet 实现了 0.92 的最佳 AUROC 和 0.94 的平均精度。此外,HPOEmb-Orphanet 的最佳 F1 分数为 0.86。定量患者相似性测量任务表明,HPOEmb-Orphanet 对 10 多种罕见疾病的相似患者实现了最高的平均检出率,并且比现有工具 HPOSim 实施的其他相似性测量表现更好,特别是对于共享共同表型较少的成对患者。定性评估表明,丰富的 HPO 嵌入通常能够以细粒度检测节点之间的关系,并且 HPOEmb-Orphanet 特别擅长关联不同疾病系统的表型。对于检测给定 PH 相关表型的相关表型特征的用例,HPOEmb-Orphanet 的表现优于其他三种 HPO 嵌入,实现了最高平均 P@5 0.81 和最高 P@10 0.79。与 HPOSim 提供的七种传统相似性测量相比,HPOEmb-Orphanet 能够检测更多相关的表型对,特别是对于不具有遗传关系的表型对。结论:根据评估结果,我们得出以下结论。首先,通过额外的非继承边,丰富的 HPO 嵌入可以检测细粒度表型节点之间的更多关联,无论它们在 HPO 图中的拓扑结构如何。其次,HPOEmb-Orphanet 不仅可以通过基于表型相似性的链接预测和患者分层来实现最佳性能,而且还能够检测到比其他嵌入和传统相似性测量更接近领域专家判断的相关表型。第三,整合异构知识资源并不一定会带来更好的检测相关表型的性能。 From a clinical perspective, in our use case study, clinical-oriented knowledge resources (e.g., Orphanet) can achieve better performance in detecting relevant phenotypic characterizations compared to biomedical-oriented knowledge resources (e.g., DECIPHER and OMIM).
Background: In precision medicine, deep phenotyping is defined as the precise and comprehensive analysis of phenotypic abnormalities, aiming to acquire a better understanding of the natural history of a disease and its genotype-phenotype associations. Detecting phenotypic relevance is an important task when translating precision medicine into clinical practice, especially for patient stratification tasks based on deep phenotyping. In our previous work, we developed node embeddings for the Human Phenotype Ontology (HPO) to assist in phenotypic relevance measurement incorporating distributed semantic representations. However, the derived HPO embeddings hold only distributed representations for IS-A relationships among nodes, hampering the ability to fully explore the graph.Methods: In this study, we developed a framework, HPO2Vec +, to enrich the produced HPO embeddings with heterogeneous knowledge resources (i.e., DECIPHER, OMIM, and Orphanet) for detecting phenotypic relevance. Specifically, we parsed disease-phenotype associations contained in these three resources to enrich non-inheritance relationships among phenotypic nodes in the HPO. To generate node embeddings for the HPO, no-de2vec was applied to perform node sampling on the enriched HPO graphs based on random walk followed by feature learning over the sampled nodes to generate enriched node embeddings. Four HPO embeddings were generated based on different graph structures, which we hereafter label as HPOEmb-Original, HPOEmb-DECIPHER, HPOEmb-OMIM, and HPOEmb-Orphanet. We evaluated the derived embeddings quantitatively through an HPO link prediction task with four edge embeddings operations and six machine learning algorithms. The resulting best embeddings were then evaluated for patient stratification of 10 rare diseases using electronic health records (EHR) collected at Mayo Clinic. We assessed our framework qualitatively by visualizing phenotypic clusters and conducting a use case study on primary hyperoxaluria (PH), a rare disease, on the task of inferring relevant phenotypes given 22 annotated PH related phenotypes.Results: The quantitative link prediction task shows that HPOEmb-Orphanet achieved an optimal AUROC of 0.92 and an average precision of 0.94. In addition, HPOEmb-Orphanet achieved an optimal F1 score of 0.86. The quantitative patient similarity measurement task indicates that HPOEmb-Orphanet achieved the highest average detection rate for similar patients over 10 rare diseases and performed better than other similarity measures implemented by an existing tool, HPOSim, especially for pairwise patients with fewer shared common phenotypes. The qualitative evaluation shows that the enriched HPO embeddings are generally able to detect relationships among nodes with fine granularity and HPOEmb-Orphanet is particularly good at associating phenotypes across different disease systems. For the use case of detecting relevant phenotypic characterizations for given PH related phenotypes, HPOEmb-Orphanet outperformed the other three HPO embeddings by achieving the highest average P@5 of 0.81 and the highest P@10 of 0.79. Compared to seven conventional similarity measurements provided by HPOSim, HPOEmb-Orphanet is able to detect more relevant phenotypic pairs, especially for pairs not in inheritance relationships.Conclusion: We drew the following conclusions based on the evaluation results. First, with additional non-inheritance edges, enriched HPO embeddings can detect more associations between fine granularity phenotypic nodes regardless of their topological structures in the HPO graph. Second, HPOEmb-Orphanet not only can achieve the optimal performance through link prediction and patient stratification based on phenotypic similarity, but is also able to detect relevant phenotypes closer to domain expert's judgments than other embeddings and conventional similarity measurements. Third, incorporating heterogeneous knowledge resources do not necessarily result in better performance for detecting relevant phenotypes. From a clinical perspective, in our use case study, clinical-oriented knowledge resources (e.g., Orphanet) can achieve better performance in detecting relevant phenotypic characterizations compared to biomedical-oriented knowledge resources (e.g., DECIPHER and OMIM).