Clusterability and Clustering of Images and Other "Real" High-Dimensional Data

Clusterability and Clustering of Images and Other "Real" High-Dimensional Data
复制标题

DOI:
10.1109/tip.2017.2789327
复制
发表时间:
2018-04-01
影响因子:
10.6
通讯作者:
Boutin, Mireille
Boutin, Mireille
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yellamraju, Tarun;Boutin, Mireille

文献摘要

被引文献

相似文献

众所周知,对高维数据集进行聚类非常困难。在本文中,我们表明,当要聚类的点对应于图像时,情况并非如此。更具体地说,图像数据集被证明具有很多结构,如此之多,因此将数据集投影到随机一维线性子空间上可能会发现图像之间的二进制分组。基于这一观察,我们提出了一种量化数据集可聚类性的方法。该方法基于数​​据投影到随机线上的聚类性(一维)度量 (S) 的概率密度。在将图像数据集的可聚类性与合成生成的聚类的可聚类性进行比较后,我们得出结论,我们在图像数据集中发现的这些有趣的结构并不符合传统意义上的聚类概念。我们的观察进一步表明,这是一种以分层方式对高维数据进行聚类的快速方法;在每个阶段,根据数据一维随机投影中发现的二元聚类,将数据分为两部分。由于大多数计算都是在一维中执行的,因此这种方法非常有效。尽管它很简单,但它总体上比现有的高维聚类方法实现了更好的聚类质量,不仅对于表示图像数据的数据集,而且对于其他真实数据集也是如此。我们的结果强调需要重新检查我们对高维聚类和真实数据集(例如图像集)的几何形状的假设。
Clustering a high-dimensional data set is known to be very difficult. In this paper, we show that this is not the case when the points to cluster correspond to images. More specifically, image data sets are shown to have a lot of structures, so much, so that projecting the set onto a random 1D linear sub-space is likely to uncover a binary grouping among the images. Based on this observation, we propose a method to quantify the clusterability of a data set. The method is based on the probability density of a measure (S) of clusterability (in 1D) of the projection of the data onto a random line. After comparing the clusterability of image datasets with that of synthetically generated clusters, we conclude that these intriguing structures we find in image datasets do not fit the notion of clusters in the traditional sense. Further suggested by our observation is a fast method for clustering high-dimensional data in a hierarchical fashion; at each stage, the data is partitioned into two based on the binary clustering found in a 1D random projection of the data. Since most of the computations are performed in 1D, this approach is extremely efficient. But despite its simplicity, it achieves overall a better quality of clustering than existing high-dimensional clustering methods, not only for datasets representing image data, but for other real data sets as well. Our results highlight the need to re-examine our assumptions about high-dimensional clustering and the geometry of real datasets such as sets of images.