TSA-tree: a wavelet-based approach to improve the efficiency of multi-level surprise and trend queries on time-series data

TSA-tree: a wavelet-based approach to improve the efficiency of multi-level surprise and trend queries on time-series data
复制标题

DOI:
10.1109/ssdm.2000.869778
复制
发表时间:
2000-07
期刊:
Proceedings. 12th International Conference on Scientific and Statistica Database Management
影响因子:
--
通讯作者:
C. Shahabi;Xiaoming Tian;Wugang Zhao
C. Shahabi;Xiaoming Tian;Wugang Zhao
中科院分区:
其他
文献类型:
--
作者:
C. Shahabi;Xiaoming Tian;Wugang Zhao

文献摘要

被引文献

相似文献

我们引入了一种新颖的基于小波的树结构,称为TSA树,它提高了对时间序列数据进行多级趋势和惊喜查询的效率。随着概念化为时间序列的科学观测数据的爆炸式增长,我们面临着有效存储、检索和分析这些数据的挑战。对此数据集的频繁查询是为了在原始时间序列中查找趋势(例如,全球变暖)或意外情况(例如,海底火山喷发)。然而,挑战在于不同的抽象级别都需要这些趋势和意外查询。为了支持这些多级趋势和意外查询,有时需要检索和处理大量原始数据。为了加快这一过程,我们利用了 TSA 树。 TSA 树的每个节点都包含不同级别的预先计算的趋势和意外情况。递归地使用小波变换来构建 TSA 节点。因此,TSA 树的每个节点都可以随时用于趋势和意外的可视化。此外,每个节点的大小明显小于原始时间序列的大小,从而导致 I/O 操作更快。然而 TSA 树的一个限制是它的大小比原始时间序列大。为了解决这个缺点,首先我们证明存储 TSA 树(OTSA 树)的最优子树所需的存储空间不超过存储原始时间序列而不丢失任何信息所需的存储空间。接下来,我们提出两种替代技术,以进一步减小 OTSA 树的大小,同时与查询原始时间序列相比保持可接受的查询精度。利用真实和合成的时间序列数据库,我们将我们的技术与一些众所周知的算法进行比较。
We introduce a novel wavelet based tree structure, termed TSA-tree, which improves the efficiency of multi-level trend and surprise queries on time sequence data. With the explosion of scientific observation data conceptualized as time sequences, we are facing the challenge of efficiently storing, retrieving and analyzing this data. Frequent queries on this data set are to find trends (e.g., global warming) or surprises (e.g., undersea volcano eruption) within the original time series. The challenge, however is that these trend and surprise queries are needed at different levels of abstractions. To support these multi-level trend and surprise queries, sometimes a huge subset of raw data needs to be retrieved and processed. To expedite this process, we utilize our TSA-tree. Each node of the TSA-tree contains pre-computed trends and surprises at different levels. A wavelet transform is used recursively to construct TSA nodes. As a result, each node of TSA tree is readily available for visualization of trends and surprises. In addition, the size of each node is significantly smaller than that of the original time series, resulting in faster I/O operations. However a limitation of TSA-tree is that its size is larger than the original time series. To address this shortcoming, first we prove that the storage space required to store the optimal subtree of TSA-tree (OTSA-tree) is no more than that required to store the original time series without losing any information. Next, we propose two alternative techniques to reduce the size of the OTSA-tree even further while maintaining an acceptable query precision as compared to querying the original time sequences. Utilizing real and synthetic time sequence databases, we compare our techniques with some well known algorithms.