Clustering multivariate time series using energy distance

Clustering multivariate time series using energy distance
复制标题

DOI:
10.1111/jtsa.12688
复制
发表时间:
2023-03
影响因子:
0.9
通讯作者:
R. Davis;Leon Fernandes;K. Fokianos
R. Davis;Leon Fernandes;K. Fokianos
中科院分区:
数学4区
文献类型:
--
作者:
R. Davis;Leon Fernandes;K. Fokianos

文献摘要

相似文献

提出了一种新的方法,用于使用Székely和Rizzo(2013)中定义的能量距离对多变量时间序列数据进行聚类。具体而言,使用能量距离统计量来形成相异性矩阵,以测量分量时间序列的有限维分布之间的分离。一旦计算成对相异性矩阵,然后应用层次聚类方法来获得树状图。这个过程是完全非参数的,因为固定分布之间的差异是直接计算的,而不需要做任何模型假设。为了证明这一过程,渐近性质的能量距离估计一般平稳和遍历时间序列。该方法说明了在模拟研究中的各种组件的时间序列,无论是线性或非线性。最后,该方法被应用到两个例子,一个涉及选定的国家的国内生产总值,另一个是美国各州的人口规模在1900年至1999年。
A novel methodology is proposed for clustering multivariate time series data using energy distance defined in Székely and Rizzo (2013). Specifically, a dissimilarity matrix is formed using the energy distance statistic to measure the separation between the finite‐dimensional distributions for the component time series. Once the pairwise dissimilarity matrix is calculated, a hierarchical clustering method is then applied to obtain the dendrogram. This procedure is completely nonparametric as the dissimilarities between stationary distributions are directly calculated without making any model assumptions. In order to justify this procedure, asymptotic properties of the energy distance estimates are derived for general stationary and ergodic time series. The method is illustrated in a simulation study for various component time series that are either linear or nonlinear. Finally, the methodology is applied to two examples; one involves the GDP of selected countries and the other is the population size of various states in the U.S.A. in the years 1900–1999.