Robust Testing in High-Dimensional Sparse Models

Robust Testing in High-Dimensional Sparse Models
复制标题

高维稀疏模型中的稳健测试

DOI:
10.48550/arxiv.2205.07488
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
C. Canonne
C. Canonne
中科院分区:
--
文献类型:
--
作者:
Anand George;C. Canonne

文献摘要

参考文献

被引文献

相似文献

我们考虑了在两种不同的观测模型下对高维稀疏信号向量的范数进行鲁棒检验的问题。在第一个模型中,我们给出$n$ i.i.d.。来自分布$\mathcal{N}\left(\theta,I_d\right)$(未知$\theta$)的样本,其中一小部分被任意损坏。在$\|\theta\|_0\le $的承诺下,我们希望正确区分对于某个输入参数$\gamma>0$,是$\|\theta\|_2=0$还是$\|\theta\|_2>\gamma$。我们表明,任何算法这个任务需要$n=\Omega\left(s\log\frac{艾德}{s}\right)$ samples,这是紧对数因子。我们还将我们的结果扩展到其他常见的稀疏性概念,即,$\|\theta\|_q\le s$的任何$0<q<2$。在我们考虑的第二个观察模型中,数据是根据稀疏线性回归模型生成的,其中协变量是i.i.d.。高斯和回归系数(信号)是已知的$s$-稀疏。这里我们也假设数据的$\n $部分是任意损坏的。我们证明了任何可靠地检验回归系数范数的算法至少需要$n=\Omega\left(\min(s\log d,{1}/{\gamma^4})\right)$ samples。我们的研究结果表明,在这两种设置下的测试的复杂性显着增加鲁棒性约束。这与最近在稳健均值检验和稳健协方差检验中观察到的结果一致。
We consider the problem of robustly testing the norm of a high-dimensional sparse signal vector under two different observation models. In the first model, we are given $n$ i.i.d. samples from the distribution $\mathcal{N}\left(\theta,I_d\right)$ (with unknown $\theta$), of which a small fraction has been arbitrarily corrupted. Under the promise that $\|\theta\|_0\le s$, we want to correctly distinguish whether $\|\theta\|_2=0$ or $\|\theta\|_2>\gamma$, for some input parameter $\gamma>0$. We show that any algorithm for this task requires $n=\Omega\left(s\log\frac{ed}{s}\right)$ samples, which is tight up to logarithmic factors. We also extend our results to other common notions of sparsity, namely, $\|\theta\|_q\le s$ for any $0<q<2$. In the second observation model that we consider, the data is generated according to a sparse linear regression model, where the covariates are i.i.d. Gaussian and the regression coefficient (signal) is known to be $s$-sparse. Here too we assume that an $\epsilon$-fraction of the data is arbitrarily corrupted. We show that any algorithm that reliably tests the norm of the regression coefficient requires at least $n=\Omega\left(\min(s\log d,{1}/{\gamma^4})\right)$ samples. Our results show that the complexity of testing in these two settings significantly increases under robustness constraints. This is in line with the recent observations made in robust mean testing and robust covariance testing.
DOI: --
发表时间: 2020-05
期刊: --
影响因子: --
作者:
Matthew Brennan;Guy Bresler
通讯作者: Matthew Brennan;Guy Bresler
鲁棒协方差检验的样本复杂性
DOI: --
发表时间: 2021
期刊: 2021
影响因子: --
作者:
Ilias Diakonikolas;Daniel M. Kane
通讯作者: Daniel M. Kane
稀疏线性回归中的全有或全无现象
DOI: --
发表时间: 2019
期刊: Proceedings of Machine Learning Research
影响因子: --
作者:
Reeves, Galen;Xu, Jiaming;Zadik, Ilias
通讯作者: Zadik, Ilias
通过迭代过滤的异常值鲁棒高维稀疏估计
DOI: --
发表时间: 2019
期刊: Advances in neural information processing systems
影响因子: --
作者:
Diakonikolas, Ilias;Kane, Daniel;Karmalkar, Sushrut;Price, Eric;Stewart, Alistair
通讯作者: Stewart, Alistair