Efficient Statistics for Sparse Graphical Models from Truncated Samples

Efficient Statistics for Sparse Graphical Models from Truncated Samples
复制标题

来自截断样本的稀疏图形模型的有效统计

DOI:
--
复制
发表时间:
2020
期刊:
International Conference on Artificial Intelligence and Statistics
影响因子:
--
通讯作者:
Ioannis Panageas
Ioannis Panageas
中科院分区:
--
文献类型:
--
作者:
Arnab Bhattacharyya;Rathin Desai;Sai Ganesh Nagarajan;Ioannis Panageas

文献摘要

参考文献

被引文献

相似文献

本文研究了截断样本的高维估计问题。我们关注两个基本的经典问题:(I)稀疏高斯图模型的推断和(Ii)稀疏线性模型的支持恢复。 (I)对于高斯图模型,假设$d$维样本${\bfx}$是由高斯$N(\u,\Sigma)$生成的,并且仅当它们属于一个子集$S\subseteq\mathbb{R}^d$时才能观察到。我们证明,使用截断的$\Mathcal{N}({\Mu},{\Sigma})$的$\tilde{O}\left(\frac{\textrm{nz}({\Sigma}^{-1})}{\epsilon^2}\right)$样本,并且可以访问$S$的成员资格预言,可以在Frobenius范数中以误差$\epsilon$来估计${\Mu}$和${\Sigma}$。假设集合$S$在未知分布下具有非平凡测度,但在其他情况下是任意的。 (Ii)对于稀疏线性回归,假设样本$({\bfx},y)$生成,其中$y={\bfx}^\top{{\omega}^*}+\mathcal{N}(0,1)$,且$({\bfx},y)$仅当$y$属于截断集$S\subseteq\mathbb{R}$时可见。我们考虑这样的情况:${\Omega}^*$是稀疏的,有一个大小为$k$的支持集。我们的主要结果是在问题维度$d$、支持度$k$、观察次数$n$以及样本的性质和截断上建立了足以恢复${\Omega}^*$的支持度的精确条件。具体地说,我们证明了在一些较温和的假设下,只需要$O(k^2\logd)$样本就可以估计$\ell_\inty$-范数中的${\Omega}^*$,直到有界误差。 对于这两个问题,我们的估计量最小化了有限总体负对数似然函数和$\ell_1$-正则化项的和。
In this paper, we study high-dimensional estimation from truncated samples. We focus on two fundamental and classical problems: (i) inference of sparse Gaussian graphical models and (ii) support recovery of sparse linear models. (i) For Gaussian graphical models, suppose $d$-dimensional samples ${\bf x}$ are generated from a Gaussian $N(\mu,\Sigma)$ and observed only if they belong to a subset $S \subseteq \mathbb{R}^d$. We show that ${\mu}$ and ${\Sigma}$ can be estimated with error $\epsilon$ in the Frobenius norm, using $\tilde{O}\left(\frac{\textrm{nz}({\Sigma}^{-1})}{\epsilon^2}\right)$ samples from a truncated $\mathcal{N}({\mu},{\Sigma})$ and having access to a membership oracle for $S$. The set $S$ is assumed to have non-trivial measure under the unknown distribution but is otherwise arbitrary. (ii) For sparse linear regression, suppose samples $({\bf x},y)$ are generated where $y = {\bf x}^\top{{\Omega}^*} + \mathcal{N}(0,1)$ and $({\bf x}, y)$ is seen only if $y$ belongs to a truncation set $S \subseteq \mathbb{R}$. We consider the case that ${\Omega}^*$ is sparse with a support set of size $k$. Our main result is to establish precise conditions on the problem dimension $d$, the support size $k$, the number of observations $n$, and properties of the samples and the truncation that are sufficient to recover the support of ${\Omega}^*$. Specifically, we show that under some mild assumptions, only $O(k^2 \log d)$ samples are needed to estimate ${\Omega}^*$ in the $\ell_\infty$-norm up to a bounded error. For both problems, our estimator minimizes the sum of the finite population negative log-likelihood function and an $\ell_1$-regularization term.
DOI: 10.1137/1.9781611975031.171
发表时间: 2017-04
期刊: ArXiv
影响因子: --
作者:
Ilias Diakonikolas;Gautam Kamath;D. Kane;Jerry Li;Ankur Moitra;Alistair Stewart
通讯作者: Ilias Diakonikolas;Gautam Kamath;D. Kane;Jerry Li;Ankur Moitra;Alistair Stewart
DOI: --
发表时间: 2017-03
期刊: --
影响因子: --
作者:
Ilias Diakonikolas;Gautam Kamath;D. Kane;Jerry Li;Ankur Moitra;Alistair Stewart
通讯作者: Ilias Diakonikolas;Gautam Kamath;D. Kane;Jerry Li;Ankur Moitra;Alistair Stewart
通过截断的样本进行高维度的高效统计
DOI: --
发表时间: 2018
期刊: Annual Symposium on Foundations of Computer Science
影响因子: --
作者:
Daskalakis, C.;Gouleakis, T.;Tzamos, C.;Zampetakis, M.
通讯作者: Zampetakis, M.
DOI: --
发表时间: 2019
期刊: Proceedings of Machine Learning Research
影响因子: --
作者:
Daskalakis, Constantinos;Gouleakis, Themis;Tzamos, Christos;Zampetakis, Manolis
通讯作者: Zampetakis, Manolis