On clustering validation techniques

On clustering validation techniques
复制标题

DOI:
10.1023/a:1012801612483
复制
发表时间:
2001-01-01
影响因子:
3.4
通讯作者:
Vazirgiannis, M
Vazirgiannis, M
中科院分区:
计算机科学3区
文献类型:
--
作者:
Halkidi, M;Batistakis, Y;Vazirgiannis, M

文献摘要

被引文献

相似文献

聚类分析旨在识别相似对象的组,因此有助于发现大型数据集中的模式分布和有趣的相关性。由于它在工程、商业和社会科学的许多应用领域中出现,因此一直是广泛研究的主题。特别是,在过去的几年中,巨大的交易和实验数据集的可用性和数据挖掘的需求不断增加,需要聚类算法的规模,可应用于不同domines.This本文介绍了聚类的基本概念,同时它调查了广泛知名的聚类算法在比较的方式。此外,它解决了一个重要的问题,聚类过程中的质量评估的聚类结果。这也与所关注的数据集的固有特征有关。聚类有效性的措施和方法在文献中的审查。此外,本文还指出了目前聚类算法中存在的问题,并展望了聚类算法的发展趋势。
Cluster analysis aims at identifying groups of similar objects and, therefore helps to discover distribution of patterns and interesting correlations in large data sets. It has been subject of wide research since it arises in many application domains in engineering, business and social sciences. Especially, in the last years the availability of huge transactional and experimental data sets and the arising requirements for data mining created needs for clustering algorithms that scale and can be applied in diverse domains.This paper introduces the fundamental concepts of clustering while it surveys the widely known clustering algorithms in a comparative way. Moreover, it addresses an important issue of clustering process regarding the quality assessment of the clustering results. This is also related to the inherent features of the data set under concern. A review of clustering validity measures and approaches available in the literature is presented. Furthermore, the paper illustrates the issues that are under-addressed by the recent algorithms and gives the trends in clustering process.