Probabilistic Sparse Subspace Clustering Using Delayed Association

Probabilistic Sparse Subspace Clustering Using Delayed Association
复制标题

DOI:
10.1109/icpr.2018.8545569
复制
发表时间:
2018-08
期刊:
2018 24th International Conference on Pattern Recognition (ICPR)
影响因子:
--
通讯作者:
Maryam Jaberi;M. Pensky;H. Foroosh
Maryam Jaberi;M. Pensky;H. Foroosh
中科院分区:
其他
文献类型:
--
作者:
Maryam Jaberi;M. Pensky;H. Foroosh

文献摘要

相似文献

发现和聚类高维数据中的子空间是机器学习的一个基本问题,在数据挖掘、计算机视觉和模式识别中有着广泛的应用。早期的方法将问题分为两个独立的阶段,即找到相似性矩阵和找到聚类。类似于最近的一些作品,我们整合这两个步骤使用联合优化方法。我们做了以下贡献:(i)我们估计每个点的集群分配的可靠性之前,分配一个点的子空间。我们将数据点分为两组“确定”和“不确定”,后一组的分配延迟,直到他们的子空间关联确定性提高。(ii)我们证明了延迟关联更适合于聚类具有模糊性的子空间,即当子空间相交或数据被离群值/噪声污染时。(iii)我们的实验表明,这种延迟的概率关联导致更准确的自我表示和最终的集群。所提出的方法具有较高的精度都只位于一个子空间的点,和那些在子空间的交集。(iv)我们表明,延迟关联导致计算成本的巨大减少,因为它允许增量谱聚类。
Discovering and clustering subspaces in high-dimensional data is a fundamental problem of machine learning with a wide range of applications in data mining, computer vision, and pattern recognition. Earlier methods divided the problem into two separate stages of finding the similarity matrix and finding clusters. Similar to some recent works, we integrate these two steps using a joint optimization approach. We make the following contributions: (i) we estimate the reliability of the cluster assignment for each point before assigning a point to a subspace. We group the data points into two groups of “certain” and “uncertain”, with the assignment of latter group delayed until their subspace association certainty improves. (ii) We demonstrate that delayed association is better suited for clustering subspaces that have ambiguities, i.e. when subspaces intersect or data are contaminated with outliers/noise. (iii) We demonstrate experimentally that such delayed probabilistic association leads to a more accurate self-representation and final clusters. The proposed method has higher accuracy both for points that exclusively lie in one subspace, and those that are on the intersection of subspaces. (iv) We show that delayed association leads to huge reduction of computational cost, since it allows for incremental spectral clustering.