Subspace Clustering through Sub-Clusters

Subspace Clustering through Sub-Clusters
复制标题

DOI:
--
复制
发表时间:
2018-11
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Weiwei Li;Jan Hannig;S. Mukherjee
Weiwei Li;Jan Hannig;S. Mukherjee
中科院分区:
其他
文献类型:
--
作者:
Weiwei Li;Jan Hannig;S. Mukherjee

文献摘要

相似文献

降维问题在现代数据分析中越来越重要。在本文中,我们考虑在一个高维空间作为一个低维子空间的联合建模的点的集合。特别是,我们提出了一个高度可扩展的基于采样的算法,集群的整个数据,通过第一个频谱聚类的一个小的随机样本,然后分类或标记其余的样本点。关键的想法是,这个随机子集借用整个数据集的信息,聚类点的问题可以用更有效和更鲁棒的“聚类子聚类”问题来代替。我们为我们的程序提供理论保证。数值结果表明,我们在准确性和速度方面优于其他最先进的子空间聚类算法。
The problem of dimension reduction is of increasing importance in modern data analysis. In this paper, we consider modeling the collection of points in a high dimensional space as a union of low dimensional subspaces. In particular we propose a highly scalable sampling based algorithm that clusters the entire data via first spectral clustering of a small random sample followed by classifying or labeling the remaining out of sample points. The key idea is that this random subset borrows information across the entire data set and that the problem of clustering points can be replaced with the more efficient and robust problem of "clustering sub-clusters". We provide theoretical guarantees for our procedure. The numerical results indicate we outperform other state-of-the-art subspace clustering algorithms with respect to accuracy and speed.