Hierarchical Clustering Based on Mutual Information

Hierarchical Clustering Based on Mutual Information
复制标题

DOI:
--
复制
发表时间:
2003-11
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Kraskov;H. Stögbauer;R. Andrzejak;P. Grassberger
A. Kraskov;H. Stögbauer;R. Andrzejak;P. Grassberger
中科院分区:
其他
文献类型:
--
作者:
A. Kraskov;H. Stögbauer;R. Andrzejak;P. Grassberger

文献摘要

被引文献

相似文献

动机:在各种生物信息应用中,集群是一个经常使用的概念。提出了一种新的数据层次聚类方法--互信息聚类(MIC)算法。它使用互信息(MI)作为相似性度量,并利用其分组特性:三个对象X、Y和Z之间的MI等于X和Y之间的MI之和,加上Z和组合对象(XY)之间的MI。结果:我们在信息理论的Shannon(概率)版本中和Kolmogorov(算法)版本中都使用了这一点,在香农(概率)版本中,“对象”是由随机样本表示的概率分布,在该版本中,“对象”是符号序列。我们将我们的方法应用于从线粒体DNA序列构建哺乳动物系统发育树,并根据独立分量分析(ICA)的输出重建胎儿心电,并将其应用于孕妇的心电。可用性:估算MI和集群的程序(概率版)可从以下http URL获得
Motivation: Clustering is a frequently used concept in variety of bioinformatical applications. We present a new method for hierarchical clustering of data called mutual information clustering (MIC) algorithm. It uses mutual information (MI) as a similarity measure and exploits its grouping property: The MI between three objects X, Y, and Z is equal to the sum of the MI between X and Y, plus the MI between Z and the combined object (XY). Results: We use this both in the Shannon (probabilistic) version of information theory, where the "objects" are probability distributions represented by random samples, and in the Kolmogorov (algorithmic) version, where the "objects" are symbol sequences. We apply our method to the construction of mammal phylogenetic trees from mitochondrial DNA sequences and we reconstruct the fetal ECG from the output of independent components analysis (ICA) applied to the ECG of a pregnant woman. Availability: The programs for estimation of MI and for clustering (probabilistic version) are available at this http URL