Location-Awareness in Time Series Compression

Location-Awareness in Time Series Compression
复制标题

时间序列压缩中的位置感知

DOI:
--
复制
发表时间:
2018
期刊:
Symposium on Advances in Databases and Information Systems
影响因子:
--
通讯作者:
D. Klabjan
D. Klabjan
中科院分区:
--
文献类型:
--
作者:
Xu Teng;Andreas Züfle;Goce Trajcevski;D. Klabjan

文献摘要

被引文献

相似文献

我们提出了关于时间序列压缩可能对相似性查询产生影响的问题的初步发现,在数据集元素伴随着附加上下文的设置中。从广义上讲,任何数据压缩方法的主要目标都是为给定的原始数据集提供更紧凑(即更小的大小)的表示。然而,正如在空间数据压缩的大量工作中所观察到的那样,“盲目地”应用特定算法可能会产生违背直觉预期的结果——例如,扭曲存在于“原始”数据[7]中的某些拓扑关系。在本研究中,我们通过定义基于肯德尔(au)的相似性扭曲度量来量化这种扭曲。我们对该度量进行了评估,并对五种最常用的时间序列压缩算法和三种最常用的时间序列相似性度量实现了相应的压缩比。我们在这里报告了我们的一些观察结果,以及对可能产生的更广泛影响和我们计划在未来应对的挑战的讨论。
We present our initial findings regarding the problem of the impact that time series compression may have on similarity-queries, in the settings in which the elements of the dataset are accompanied with additional contexts. Broadly, the main objective of any data compression approach is to provide a more compact (i.e., smaller size) representation of a given original dataset. However, as has been observed in the large body of works on compression of spatial data, applying a particular algorithm “blindly” may yield outcomes that defy the intuitive expectations – e.g., distorting certain topological relationships that exist in the “raw” data [7]. In this study, we quantify this distortion by defining a measure of similarity distortion based on Kendall’s ( au ). We evaluate this measure, and the correspondingly achieved compression ratio for the five most commonly used time series compression algorithms and the three most common time series similarity measures. We report some of our observations here, along with the discussion of the possible broader impacts and the challenges that we plan to address in the future.