Independent Component Analysis Based Seeding Method for K-Means Clustering

Independent Component Analysis Based Seeding Method for K-Means Clustering
复制标题

DOI:
10.1109/wi-iat.2011.29
复制
发表时间:
2011-08
期刊:
2011 IEEE/WIC/ACM International Conferences on Web Intelligence and Intelligent Agent Technology
影响因子:
--
通讯作者:
T. Onoda;Miho Sakai;S. Yamada
T. Onoda;Miho Sakai;S. Yamada
中科院分区:
其他
文献类型:
--
作者:
T. Onoda;Miho Sakai;S. Yamada

文献摘要

相似文献

k-means聚类方法是一种广泛使用的Web聚类技术,因为它的简单性和速度。然而,聚类结果在很大程度上取决于所选择的初始聚类中心,这是均匀随机选择的数据点。提出了一种基于独立成分分析的种子算法。我们评估我们提出的方法的性能,并通过使用基准数据集与其他播种方法进行比较。我们将我们提出的方法应用于Web语料库,这是由ODP提供的。实验表明,该方法的归一化互信息优于k-means聚类方法和k-means ++聚类方法的归一化互信息。因此,该方法对Web语料库具有一定的实用价值。
The k-means clustering method is a widely used clustering technique for the Web because of its simplicity and speed. However, the clustering result depends heavily on the chosen initial clustering centers, which are chosen uniformly at random from the data points. We propose a seeding method based on the independent component analysis for the k-means clustering method. We evaluate the performance of our proposed method and compare it with other seeding methods by using benchmark datasets. We applied our proposed method to a Web corpus, which is provided by ODP. The experiments show that the normalized mutual information of our proposed method is better than the normalized mutual information of k-means clustering method and k-means++ clustering method. Therefore, the proposed method is useful for Web corpus.