Grouping Methods for Pattern Matching over Probabilistic Data Streams

Grouping Methods for Pattern Matching over Probabilistic Data Streams
复制标题

DOI:
10.1587/transinf.2016dap0014
复制
发表时间:
2017-04
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Kento Sugiura;Y. Ishikawa;Yuya Sasaki
Kento Sugiura;Y. Ishikawa;Yuya Sasaki
中科院分区:
其他
文献类型:
--
作者:
Kento Sugiura;Y. Ishikawa;Yuya Sasaki

文献摘要

相似文献

随着传感器和机器学习技术的发展,从概率数据流中检测模式变得越来越重要。本文主要研究基于模式匹配的复杂事件处理。将模式匹配应用于概率数据流时,由于数据的不确定性,可能会在同一时间间隔内检测到大量的匹配。尽管现有的方法对这些匹配进行了区分,但是当其中一些匹配对应于在时间间隔内发生的真实事件时,它们可能会得出不合适的结果。因此,我们提出了两种匹配分组方法。我们的方法输出指示在给定时间间隔内发生的复杂事件的组。本文首先描述了基于时间重叠的分组的定义,并提出了两种分组算法,引入了完全重叠和单重叠的概念。然后,我们提出了一种高效的方法,通过使用由查询模式生成的确定性有限自动机来计算组的出现概率。最后,我们通过将我们的方法应用于真实和合成数据集来经验地评估它们的有效性。
SUMMARY As the development of sensor and machine learning technologies has progressed, it has become increasingly important to detect patterns from probabilistic data streams . In this paper, we focus on complex event processing based on pattern matching . When we apply pattern matching to probabilistic data streams, numerous matches may be detected at the same time interval because of the uncertainty of data. Although existing methods distinguish between such matches, they may derive inappropriate results when some of the matches correspond to the real-world event that has occurred during the time interval. Thus, we propose two grouping methods for matches. Our methods output groups that indicate the occurrence of complex events during the given time intervals. In this paper, first we describe the definition of groups based on temporal overlap, and propose two grouping algorithms, introducing the notions of complete overlap and single overlap . Then, we propose an e ffi cient approach for calculating the occurrence probabilities of groups by using deterministic finite automata that are generated from the query patterns. Finally, we empirically evaluate the e ff ectiveness of our methods by applying them to real and synthetic datasets.