High-Dimensional Graphical Model Selection Using $\ell_1$-Regularized Logistic Regression

High-Dimensional Graphical Model Selection Using $\ell_1$-Regularized Logistic Regression
复制标题

DOI:
--
复制
发表时间:
2008-04
期刊:
arXiv: Statistics Theory
影响因子:
--
通讯作者:
Pradeep Ravikumar;M. Wainwright;J. Lafferty
Pradeep Ravikumar;M. Wainwright;J. Lafferty
中科院分区:
其他
文献类型:
--
作者:
Pradeep Ravikumar;M. Wainwright;J. Lafferty

文献摘要

被引文献

相似文献

我们考虑的问题,估计与离散马尔可夫随机场的图结构。我们描述了一种方法的基础上$\ell_1$-正则化逻辑回归,其中任何给定的节点的邻域估计进行逻辑回归受$\ell_1$-约束。我们的框架适用于高维设置,其中节点的数量$p$和最大邻域大小$d$都允许增长作为观测数量$n$的函数。我们的主要结果提供了充分条件的三元组$(n,p,d)$的方法,以成功地在一致的估计在图中的每个节点的邻域同时。在对总体Fisher信息矩阵的一定假设下,我们证明了当样本容量为n = \Omega(d^3 \log p)时,可以得到一致的邻域选择,并且对于某个常数$C$,误差衰减为$\order(\exp(-Cn/d^3))$.如果这些相同的假设直接施加在样本矩阵上,我们证明了$n = \Omega(d^2 \log p)$ samples是足够的。
We consider the problem of estimating the graph structure associated with a discrete Markov random field. We describe a method based on $\ell_1$-regularized logistic regression, in which the neighborhood of any given node is estimated by performing logistic regression subject to an $\ell_1$-constraint. Our framework applies to the high-dimensional setting, in which both the number of nodes $p$ and maximum neighborhood sizes $d$ are allowed to grow as a function of the number of observations $n$. Our main results provide sufficient conditions on the triple $(n, p, d)$ for the method to succeed in consistently estimating the neighborhood of every node in the graph simultaneously. Under certain assumptions on the population Fisher information matrix, we prove that consistent neighborhood selection can be obtained for sample sizes $n = \Omega(d^3 \log p)$, with the error decaying as $\order(\exp(-C n/d^3))$ for some constant $C$. If these same assumptions are imposed directly on the sample matrices, we show that $n = \Omega(d^2 \log p)$ samples are sufficient.