A Better Alternative to Piecewise Linear Time Series Segmentation

A Better Alternative to Piecewise Linear Time Series Segmentation
复制标题

DOI:
10.1137/1.9781611972771.59
复制
发表时间:
2006-05
期刊:
ArXiv
影响因子:
--
通讯作者:
D. Lemire
D. Lemire
中科院分区:
其他
文献类型:
--
作者:
D. Lemire

文献摘要

被引文献

相似文献

时间序列很难监控、总结和预测。分割将时间序列组织成具有统一特征(平坦度、线性度、模态性、单调性等)的几个区间。对于可扩展性,我们需要快速的线性时间算法。流行的分段线性模型可以确定数据上升或下降的位置以及速度。不幸的是,当数据不遵循线性模型时,局部斜率的计算会产生过拟合。我们提出了一个自适应的时间序列模型,其中每个区间的多项式度变化(常数,线性等)。给定一些回归量,每个区间的成本是它的多项式度:常数区间花费1个回归量,线性区间花费2个回归量,以此类推。我们的目标是最小化给定模型复杂度的欧几里得(l_2)误差。实验上,我们研究了区间可以是常数或线性的模型。在合成随机漫步、历史股票市场价格和心电图上,自适应模型提供了比分段线性模型更准确的分割,而不会增加交叉验证误差或运行时间,同时为应用程序提供了更丰富的词汇表。讨论了实现问题,如数值稳定性和实际性能。
Time series are difficult to monitor, summarize and predict. Segmentation organizes time series into few intervals having uniform characteristics (flatness, linearity, modality, monotonicity and so on). For scalability, we require fast linear time algorithms. The popular piecewise linear model can determine where the data goes up or down and at what rate. Unfortunately, when the data does not follow a linear model, the computation of the local slope creates overfitting. We propose an adaptive time series model where the polynomial degree of each interval vary (constant, linear and so on). Given a number of regressors, the cost of each interval is its polynomial degree: constant intervals cost 1 regressor, linear intervals cost 2 regressors, and so on. Our goal is to minimize the Euclidean (l_2) error for a given model complexity. Experimentally, we investigate the model where intervals can be either constant or linear. Over synthetic random walks, historical stock market prices, and electrocardiograms, the adaptive model provides a more accurate segmentation than the piecewise linear model without increasing the cross-validation error or the running time, while providing a richer vocabulary to applications. Implementation issues, such as numerical stability and real-world performance, are discussed.