Fast and Scalable Algorithm for Detection of Structural Breaks in Big VAR Models

Fast and Scalable Algorithm for Detection of Structural Breaks in Big VAR Models
复制标题

DOI:
10.1080/10618600.2021.1950005
复制
发表时间:
2021-08
影响因子:
2.4
通讯作者:
Abolfazl Safikhani;Yue Bai;G. Michailidis
Abolfazl Safikhani;Yue Bai;G. Michailidis
中科院分区:
数学2区
文献类型:
--
作者:
Abolfazl Safikhani;Yue Bai;G. Michailidis

文献摘要

相似文献

摘要许多真实的时间序列数据都表现出随时间的结构变化。一个流行的模型来捕捉它们的时间依赖性是向量自回归(VAR),它可以通过时间演变的转移矩阵来适应结构变化。然后,问题变成估计结构突变点的(未知)数量以及VAR模型参数。在存在非常大的数据集的情况下,出现了另一个挑战,即如何以计算有效的方式实现这两个目标。在这篇文章中,我们提出了一种新的程序,它利用块分割方案(BSS),减少了模型参数的数量估计通过正则化最小二乘准则。具体而言,BSS检查适当定义的可用数据块,当与融合的基于套索的估计标准相结合时,导致显着的计算增益,而不影响识别结构断裂的数量和位置的统计准确性。该过程还与新的局部和穷举搜索步骤相结合,以一致地估计断点的数量和相对位置。该过程可扩展到大的高维时间序列数据集,其计算复杂性可以实现,其中n是时间序列的长度(样本大小),与需要步骤的穷举过程相比。大量的数值模拟工作的合成数据支持的理论研究结果,并说明了有吸引力的性能的程序。最后,一个神经科学数据集的应用程序展示了其在应用程序中的有用性。本文的补充文件可在线获得。
Abstract Many real time series datasets exhibit structural changes over time. A popular model for capturing their temporal dependence is that of vector autoregressions (VAR), which can accommodate structural changes through time evolving transition matrices. The problem then becomes to both estimate the (unknown) number of structural break points, together with the VAR model parameters. An additional challenge emerges in the presence of very large datasets, namely on how to accomplish these two objectives in a computational efficient manner. In this article, we propose a novel procedure which leverages a block segmentation scheme (BSS) that reduces the number of model parameters to be estimated through a regularized least-square criterion. Specifically, BSS examines appropriately defined blocks of the available data, which when combined with a fused lasso-based estimation criterion, leads to significant computational gains without compromising on the statistical accuracy in identifying the number and location of the structural breaks. This procedure is further coupled with new local and exhaustive search steps to consistently estimate the number and relative location of the break points. The procedure is scalable to big high-dimensional time series datasets with a computational complexity that can achieve , where n is the length of the time series (sample size), compared to an exhaustive procedure that requires steps. Extensive numerical work on synthetic data supports the theoretical findings and illustrates the attractive properties of the procedure. Finally, an application to a neuroscience dataset exhibits its usefulness in applications. Supplementary files for this article are available online.