Spherical k-Means++ Clustering

Spherical k-Means++ Clustering
复制标题

DOI:
10.1007/978-3-319-23240-9_9
复制
发表时间:
2015-09
期刊:
--
影响因子:
--
通讯作者:
Y. Endo;S. Miyamoto
Y. Endo;S. Miyamoto
中科院分区:
其他
文献类型:
--
作者:
Y. Endo;S. Miyamoto

文献摘要

被引文献

相似文献

k-means clustering(KM)算法,又称硬均值聚类(HCM)算法,是一种非常强大的聚类算法[1,2],但它存在严重的初值依赖性问题。为了降低这种依赖性,亚瑟和瓦西维茨基在2007年[3]提出了一种k-means ++聚类算法(KM++)。顺便说一句,在很多情况下,每个对象都分配在一个单位球体上,例如文本聚类。Dhillon和Modha在2007年[4]提出了原始的sphericalk-means聚类算法来对这些对象进行分类,Honik,Kober和Buchta在2012年[5]提出了新的sphericalk-means聚类(SKM)算法。然而,这两种算法也有同样的问题,初始值依赖KM。为此,本文讨论了以下几个问题:(1)将SKM的相异度推广到满足三角不等式;(2)提出了一种适用于该问题的SKM++算法。本文表明,SKM++的有效性在理论上是有保证的。
k-means clustering (KM) algorithm, also called hardc-means clustering (HCM) algorithm, is a very powerful clustering algorithm [1, 2], but it has a serious problem of strong initial value dependence. To decrease the dependence, Arthur and Vassilvitskii proposed an algorithm ofk-means++ clustering (KM++) algorithm on 2007 [3]. By the way, there are many case that each object is allocated on an unit sphere, e.g. text clustering. Dhillon and Modha proposed the primitive sphericalk-means clustering algorithm to classify such objects on 2007 [4] and Honik, Kober, and Buchta proposed new sphericalk-means clustering (SKM) algorithm on 2012 [5]. However, both of the algorithms also have the same problem of initial value dependence as KM. Therefore, the paper discuss the following points: (1) the dissimilarity of SKM is extended to satisfy the triangle inequality, and (2) sphericalk-means++ clustering (SKM++) algorithm which works well for the problem is proposed. The paper shows that the effectiveness of SKM++ is theoretically guaranteed.