Measuring Disease Similarity Based on Multiple Heterogeneous Disease Information Networks

Measuring Disease Similarity Based on Multiple Heterogeneous Disease Information Networks
复制标题

DOI:
10.1109/bibm47256.2019.8983141
复制
发表时间:
2019-11
期刊:
2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)
影响因子:
--
通讯作者:
Ling Tian;Jianliang Gao;Jianxin Wang;Ying Wang;Bo Song;Xiaohua Hu
Ling Tian;Jianliang Gao;Jianxin Wang;Ying Wang;Bo Song;Xiaohua Hu
中科院分区:
其他
文献类型:
--
作者:
Ling Tian;Jianliang Gao;Jianxin Wang;Ying Wang;Bo Song;Xiaohua Hu

文献摘要

相似文献

量化疾病之间的相似性目前在生物学和医学中发挥着重要作用,它为发现相似疾病提供了可靠的参考信息。以往的疾病间相似度计算方法要么使用单源数据,要么没有充分利用多源数据。在这项研究中,我们提出了一种利用多个异质疾病信息网络来衡量疾病相似度的方法。首先,将多个与疾病相关的数据源描述为包括疾病、途径、化学物质等各种对象的异质疾病信息网络。然后,对这些异质疾病信息网络进行顶点过滤,得到相应的子图。使用动态时间规整(DTW)算法和元路径方法分别计算了这些异构子图的拓扑分和语义分。通过这种方式,我们将多个异质疾病网络转化为一个边缘具有不同权值的同质疾病网络。最后,根据权值嵌入疾病节点,利用这些$n维向量计算疾病之间的相似度。基于Benchmark Set的实验充分证明了该方法在多源数据相似性度量中的有效性。
Quantifying the similarities between diseases is now playing an important role in biology and medicine, which provides reliable reference information in finding similar diseases. Most of the previous methods for similarity calculation between diseases either use a single-source data or do not fully utilize multi-sources data. In this study, we propose an approach to measure disease similarity by utilizing multiple heterogeneous disease information networks. Firstly, multiple disease-related data sources are formulated as heterogeneous disease information networks which include various types of objects such as disease, pathway, and chemicals. Then, the corresponding subgraphs of these heterogeneous disease information networks are obtained by filtering vertices. Topological scores and semantics scores are calculated in these heterogenous subgraphs using Dynamic Time Warping (DTW) algorithm and meta path method respectively. In this way, we transform multiple heterogeneous disease networks to a homogeneous disease network with different weights on the edges. Finally, the disease nodes can be embedded according to the weights and the similarity between diseases can then be calculated using these $n$-dimensional vectors. Experiments based on benchmark set fully demonstrate the effectiveness of our method in measuring the similarity of diseases through multi-sources data.