Sketched Subspace Clustering

Sketched Subspace Clustering
复制标题

DOI:
10.1109/tsp.2017.2781649
复制
发表时间:
2017-07
影响因子:
5.4
通讯作者:
Panagiotis A. Traganitis;G. Giannakis
Panagiotis A. Traganitis;G. Giannakis
中科院分区:
工程技术1区
文献类型:
--
作者:
Panagiotis A. Traganitis;G. Giannakis

文献摘要

被引文献

相似文献

每天产生和传播的大量数据在处理方面提出了独特的挑战。聚类,即在不存在地面实况标签的情况下对数据进行分组,是从数据中得出推论的重要工具。子空间聚类(SC)是一个相对较新的方法,能够成功地分类非线性可分离的数据在众多的设置。尽管SC方法具有很高的聚类精度,但在处理大量高维数据时,其计算复杂度非常高。受随机草图降维方法的启发,本文介绍了一种随机方案SC,称为Sketch-SC,适合于大量的高维数据。草图-SC加速计算繁重的部分国家的最先进的SC的方法,通过压缩的数据矩阵在两个维度上使用随机投影,从而实现快速,准确的大规模SC。性能分析以及广泛的数值测试真实的数据证实了潜力的草图-SC和其竞争力的性能相对于国家的最先进的可扩展的SC的方法。
The immense amount of daily generated and communicated data presents unique challenges in their processing. Clustering, the grouping of data without the presence of ground-truth labels, is an important tool for drawing inferences from data. Subspace clustering (SC) is a relatively recent method that is able to successfully classify nonlinearly separable data in a multitude of settings. In spite of their high clustering accuracy, SC methods incur prohibitively high computational complexity when processing large volumes of high-dimensional data. Inspired by random sketching approaches for dimensionality reduction, the present paper introduces a randomized scheme for SC, termed Sketch-SC, tailored for large volumes of high-dimensional data. Sketch-SC accelerates the computationally heavy parts of state-of-the-art SC approaches by compressing the data matrix across both dimensions using random projections, thus enabling fast and accurate large-scale SC. Performance analysis as well as extensive numerical tests on real data corroborate the potential of Sketch-SC and its competitive performance relative to state-of-the-art scalable SC approaches.