Low Rank and Structured Modeling of High-Dimensional Vector Autoregressions

Low Rank and Structured Modeling of High-Dimensional Vector Autoregressions
复制标题

DOI:
10.1109/tsp.2018.2887401
复制
发表时间:
2019-03-01
影响因子:
5.4
通讯作者:
Michailidis, George
Michailidis, George
中科院分区:
工程技术1区
文献类型:
--
作者:
Basu, Sumanta;Li, Xianqi;Michailidis, George

文献摘要

被引文献

相似文献

高维时间序列数据的网络建模是一项关键的学习任务,因为它在宏观经济学、金融和神经科学等许多应用领域都有广泛的应用。虽然基于向量自回归模型(VAR)的稀疏建模问题已经在文献中得到了深入的研究,但涉及低秩和组稀疏成分的更复杂的网络结构尽管存在于数据中,但受到的关注却相当少。未能解释低秩结构导致观察到的时间序列之间存在虚假的连通性,这可能导致从业者对相关的科学或政策问题得出错误的结论。为了在考虑潜在效应后准确地估计格兰杰因果相互作用网络,我们引入了一种新的方法来估计低秩和结构化的高维稀疏VAR模型。我们引入了一个包含核规范和套索(或群套索)惩罚的正则化框架。然后,我们建立了低秩和结构化稀疏分量估计错误率的非渐近概率上界。我们还介绍了一种快速估计算法,最后通过合成数据和实际数据的数值实验证明了所提出的建模框架在标准稀疏VAR估计上的性能。
Network modeling of high-dimensional time series data is a key learning task due to its widespread use in a number of application areas, including macroeconomics, finance, and neuroscience. While the problem of sparse modeling based on vector autoregressive models (VAR) has been investigated in depth in the literature, more complex network structures that involve low rank and group sparse components have received considerably less attention, despite their presence in data. Failure to account for low-rank structures results in spurious connectivity among the observed time series, which may lead practitioners to draw incorrect conclusions about pertinent scientific or policy questions. In order to accurately estimate a network of Granger causal interactions after accounting for latent effects, we introduce a novel approach for estimating low-rank and structured sparse high-dimensional VAR models. We introduce a regularized framework involving a combination of nuclear norm and lasso (or group lasso) penalties. Subsequently, we establish nonasymptotic probabilistic upper bounds on the estimation error rates of the low-rank and the structured sparse components. We also introduce a fast estimation algorithm and finally demonstrate the performance of theproposed modeling framework over standard sparse VAR estimates through numerical experiments on synthetic and real data.