Initialization of cluster refinement algorithms: a review and comparative study

Initialization of cluster refinement algorithms: a review and comparative study
复制标题

DOI:
10.1109/ijcnn.2004.1379917
复制
发表时间:
2004-07
期刊:
2004 IEEE International Joint Conference on Neural Networks (IEEE Cat. No.04CH37541)
影响因子:
--
通讯作者:
Ji He;Man Lan;C. Tan;S. Sung;H. Low
Ji He;Man Lan;C. Tan;S. Sung;H. Low
中科院分区:
其他
文献类型:
--
作者:
Ji He;Man Lan;C. Tan;S. Sung;H. Low

文献摘要

被引文献

相似文献

各种迭代求精聚类方法依赖于模型的初始状态,并且只能得到它们的局部最优解之一。由于识别全局最优解的任务是NP-Hard问题,对子优化问题的初始化方法的研究具有重要的价值。本文对文献中的各种聚类初始化方法进行了综述,将其分为三大类,即随机抽样法、距离优化法和密度估计法。此外,使用一组量化指标,我们在许多合成和真实数据集上评估了它们的性能。我们的受控基准确定了两种距离优化方法,即SCS和KKZ,作为k-Means学习特征的补充,从而在输出解决方案中实现更好的聚类分离。
Various iterative refinement clustering methods are dependent on the initial state of the model and are capable of obtaining one of their local optima only. Since the task of identifying the global optimization is NP-hard, the study of the initialization method towards a sub-optimization is of great value. This paper reviews the various cluster initialization methods in the literature by categorizing them into three major families, namely random sampling methods, distance optimization methods, and density estimation methods. In addition, using a set of quantitative measures, we assess their performance on a number of synthetic and real-life data sets. Our controlled benchmark identifies two distance optimization methods, namely SCS and KKZ, as complements of the k-means learning characteristics towards a better cluster separation in the output solution.