Data clustering: 50 years beyond K-means

Data clustering: 50 years beyond K-means
复制标题

DOI:
10.1016/j.patrec.2009.09.011
复制
发表时间:
2010-06-01
影响因子:
5.1
通讯作者:
Jain, Anil K.
Jain, Anil K.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Jain, Anil K.

文献摘要

被引文献

相似文献

将数据组织成合理的分组是理解和学习的最基本模式之一。作为一个例子,一个常见的科学分类方案将生物体放入一个分类系统:域,界,门,类等。聚类分析是根据测量或感知的内在特征或相似性对对象进行分组或聚类的方法和算法的正式研究。聚类分析不使用用先前标识符标记对象的类别标签,即,类标签。类别信息的缺失将数据聚类(无监督学习)与分类或判别分析(监督学习)区分开来。聚类的目的是发现数据中的结构,因此本质上是探索性的。聚类在各种科学领域有着悠久而丰富的历史。最流行和最简单的聚类算法之一,K-means,于1955年首次发表。尽管K-means在50多年前被提出,并且从那时起已经发布了数千种聚类算法,但K-means仍然被广泛使用。这说明了设计通用聚类算法的困难和聚类的不适定问题。我们提供了一个简要的概述聚类,总结著名的聚类方法,讨论了主要的挑战和关键问题,设计聚类算法,并指出一些新兴的和有用的研究方向,包括半监督聚类,集成聚类,同时在数据聚类的特征选择,和大规模数据聚类。(C)2009 Elsevier B. V.保留所有权利。
Organizing data into sensible groupings is one of the most fundamental modes of understanding and learning. As an example, a common scheme of scientific classification puts organisms into a system of ranked taxa: domain, kingdom, phylum, class, etc. Cluster analysis is the formal study of methods and algorithms for grouping, or clustering, objects according to measured or perceived intrinsic characteristics or similarity. Cluster analysis does not use category labels that tag objects with prior identifiers, i.e., class labels. The absence of category information distinguishes data clustering (unsupervised learning) from classification or discriminant analysis (supervised learning). The aim of clustering is to find structure in data and is therefore exploratory in nature. Clustering has a long and rich history in a variety of scientific fields. One of the most popular and simple clustering algorithms, K-means, was first published in 1955. In spite of the fact that K-means was proposed over 50 years ago and thousands of clustering algorithms have been published since then, K-means is still widely used. This speaks to the difficulty in designing a general purpose clustering algorithm and the ill-posed problem of clustering. We provide a brief overview of clustering, summarize well known clustering methods, discuss the major challenges and key issues in designing clustering algorithms, and point out some of the emerging and useful research directions, including semi-supervised clustering, ensemble clustering, simultaneous feature selection during data clustering, and large scale data clustering. (C) 2009 Elsevier B.V. All rights reserved.