Similarity-Based Clustering of Sequences Using Hidden Markov Models

Similarity-Based Clustering of Sequences Using Hidden Markov Models
复制标题

DOI:
10.1007/3-540-45065-3_8
复制
发表时间:
2003-07
期刊:
Proceedings. IEEE Computer Society Bioinformatics Conference
影响因子:
--
通讯作者:
M. Bicego;Vittorio Murino;Mário A. T. Figueiredo
M. Bicego;Vittorio Murino;Mário A. T. Figueiredo
中科院分区:
其他
文献类型:
--
作者:
M. Bicego;Vittorio Murino;Mário A. T. Figueiredo

文献摘要

被引文献

相似文献

隐马尔可夫模型构成了序列数据建模的广泛使用的工具;然而,它们在集群环境中的使用还没有得到很好的研究。本文借鉴最近在有监督学习环境中引入的基于相似性的方法,提出了一种新的基于HMM的序列数据聚类方法。通过这种方法,建立了一个新的表示空间,其中每个对象都由其相对于预定的一组其他对象的相似性向量来描述。这些相似性是使用隐马尔可夫模型确定的。然后在这样的空间中执行集群。通过这种方式,序列聚类的困难问题因此被转移到更易管理的形式,即点(特征向量)的聚类。对合成数据和真实数据的实验评估表明,该方法的性能明显优于标准的HMM聚类算法。
Hidden Markov models constitute a widely employed tool for sequential data modelling; nevertheless, their use in the clustering context has been poorly investigated. In this paper a novel scheme for HMM-based sequential data clustering is proposed, inspired on the similarity-based paradigm recently introduced in the supervised learning context. With this approach, a new representation space is built, in which each object is described by the vector of its similarities with respect to a predeterminate set of other objects. These similarities are determined using hidden Markov models. Clustering is then performed in such a space. By way of this, the difficult problem of clustering of sequences is thus transposed to a more manageable format, the clustering of points (vectors of features). Experimental evaluation on synthetic and real data shows that the proposed approach largely outperforms standard HMM clustering schemes.