Capturing the Denoising Effect of PCA via Compression Ratio

Capturing the Denoising Effect of PCA via Compression Ratio
复制标题

通过压缩比捕捉 PCA 的去噪效果

DOI:
--
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Jiapeng Zhang
Jiapeng Zhang
中科院分区:
--
文献类型:
--
作者:
Chandra Sekhar Mukherjee;Nikhil Doerkar;Jiapeng Zhang

文献摘要

参考文献

被引文献

相似文献

主成分分析(PCA)是机器学习中最基本的工具之一,广泛用作降维和去噪工具。在后一种设置中,虽然PCA已知在子空间恢复方面是有效的,并且被证明在某些特定设置中有助于聚类算法,但其对噪声数据的改善通常仍然没有很好地量化。在本文中,我们提出了一种新的度量称为压缩比,以捕获PCA对高维噪声数据的影响。我们发现,对于具有底层社区结构的数据,PCA显着降低了属于同一社区的数据点的距离,同时相对温和地降低了社区间的距离。我们通过理论证明和真实数据的实验来解释这种现象。基于这个新的度量,我们设计了一个简单的算法,可以用来检测离群值。粗略地说,我们认为具有较低压缩比方差的点不与其他点共享公共信号(因此可以被认为是离群值)。我们为这个简单的离群值检测算法提供了理论依据,并使用模拟来证明我们的方法与流行的离群值检测工具具有竞争力。最后,我们对真实世界的高维噪声数据(单细胞RNA-seq)进行了实验,以证明通过我们的离群点检测方法从这些数据集中删除点可以提高聚类算法的准确性。在这项任务中,我们的方法与流行的离群值检测工具非常有竞争力。
Principal component analysis (PCA) is one of the most fundamental tools in machine learning with broad use as a dimensionality reduction and denoising tool. In the later setting, while PCA is known to be effective at subspace recovery and is proven to aid clustering algorithms in some specific settings, its improvement of noisy data is still not well quantified in general. In this paper, we propose a novel metric called emph{compression ratio} to capture the effect of PCA on high-dimensional noisy data. We show that, for data with emph{underlying community structure}, PCA significantly reduces the distance of data points belonging to the same community while reducing inter-community distance relatively mildly. We explain this phenomenon through both theoretical proofs and experiments on real-world data. Building on this new metric, we design a straightforward algorithm that could be used to detect outliers. Roughly speaking, we argue that points that have a emph{lower variance of compression ratio} do not share a emph{common signal} with others (hence could be considered outliers). We provide theoretical justification for this simple outlier detection algorithm and use simulations to demonstrate that our method is competitive with popular outlier detection tools. Finally, we run experiments on real-world high-dimension noisy data (single-cell RNA-seq) to show that removing points from these datasets via our outlier detection method improves the accuracy of clustering algorithms. Our method is very competitive with popular outlier detection tools in this task.
论 SVD 在随机块模型中的威力
DOI: --
发表时间: 2023
期刊: NeurIPS 2023
影响因子: --
作者:
Mao, Xinyu Mao;Zhang Jiapeng
通讯作者: Zhang Jiapeng