Simultaneous Change Point Inference and Structure Recovery for High Dimensional Gaussian Graphical Models

Simultaneous Change Point Inference and Structure Recovery for High Dimensional Gaussian Graphical Models
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
B. Liu
B. Liu
中科院分区:
其他
文献类型:
--
作者:
B. Liu

文献摘要

被引文献

相似文献

在这篇文章中,我们研究了高维高斯图模型中可能发生突变的同时的变点推断和结构恢复问题。特别地,在邻域选择的启发下,我们将一个阈值变量和一个未知的阈值参数结合到一个联合稀疏回归模型中,该模型将p‘1-正则化的节点回归问题结合在一起。同时得到了精度矩阵的变点估计量和相应的估计系数。在此基础上,引入分类器来判别是否存在变化点。为了正确恢复图像的图形结构,提出了一种数据驱动的阈值分割方法。理论上,在一定的稀疏性条件和正则性假设下,我们的方法可以正确地选择同质或异质模型,并具有较高的精度。此外,在后一种情况下,通过允许节点数量远远大于样本大小,我们建立了变点估计量的估计相合性。此外,在高斯图形模型的结构恢复方面,提出的阈值方法实现了模型选择的一致性,并控制了误报的数量。通过大量的数值研究,证明了所提方法的有效性。最后,我们将所提出的方法应用于S&P500数据集,以显示其经验有效性。
In this article, we investigate the problem of simultaneous change point inference and structure recovery in the context of high dimensional Gaussian graphical models with possible abrupt changes. In particular, motivated by neighborhood selection, we incorporate a threshold variable and an unknown threshold parameter into a joint sparse regression model which combines p `1-regularized node-wise regression problems together. The change point estimator and the corresponding estimated coefficients of precision matrices are obtained together. Based on that, a classifier is introduced to distinguish whether a change point exists. To recover the graphical structure correctly, a data-driven thresholding procedure is proposed. In theory, under some sparsity conditions and regularity assumptions, our method can correctly choose a homogeneous or heterogeneous model with high accuracy. Furthermore, in the latter case with a change point, we establish estimation consistency of the change point estimator, by allowing the number of nodes being much larger than the sample size. Moreover, it is shown that, in terms of structure recovery of Gaussian graphical models, the proposed thresholding procedure achieves model selection consistency and controls the number of false positives. The validity of our proposed method is justified via extensive numerical studies. Finally, we apply our proposed method to the S&P 500 dataset to show its empirical usefulness.