Optimal string clustering based on a Laplace-like mixture and EM algorithm on a set of strings

Optimal string clustering based on a Laplace-like mixture and EM algorithm on a set of strings
复制标题

DOI:
10.1016/j.jcss.2019.07.003
复制
发表时间:
2014-11
期刊:
J. Comput. Syst. Sci.
影响因子:
--
通讯作者:
H. Koyano;M. Hayashida;T. Akutsu
H. Koyano;M. Hayashida;T. Akutsu
中科院分区:
其他
文献类型:
--
作者:
H. Koyano;M. Hayashida;T. Akutsu

文献摘要

相似文献

在这项研究中,我们解决了在无监督的方式聚类字符串数据的混合模型和EM算法的基础上,在我们以前的研究中开发的字符串的拓扑幺半群的概率理论的理论。我们开始引入一组字符串的参数概率分布,它具有字符串的位置和色散参数以及正的真实的数。我们开发了一个迭代算法估计的混合模型的参数的分布介绍,并证明我们的算法收敛到EM算法,这不能明确地写这个混合模型,概率为1,并强烈一致估计其参数作为观察字符串和迭代次数的增加。最后,我们推导出一个无监督字符串聚类的过程,在这个意义上,使正确分类的后验概率最大化是渐近最优的。
In this study, we address the problem of clustering string data in an unsupervised manner by developing a theory of a mixture model and an EM algorithm for strings based on probability theory on a topological monoid of strings developed in our previous studies. We begin with introducing a parametric probability distribution on a set of strings, which has location and dispersion parameters of a string and positive real number. We develop an iteration algorithm for estimating the parameters of the mixture model of the distributions introduced and demonstrate that our algorithm converges to the EM algorithm, which cannot be explicitly written for this mixture model, with probability one and strongly consistently estimates its parameters as the numbers of observed strings and iterations increase. We finally derive a procedure for unsupervised string clustering that is asymptotically optimal in the sense that the posterior probability of making correct classifications is maximized.