Automatic Sequential Pattern Mining in Data Streams

Automatic Sequential Pattern Mining in Data Streams
复制标题

DOI:
10.1145/3357384.3358002
复制
发表时间:
2019-11
期刊:
Proceedings of the 28th ACM International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Kouki Kawabata;Yasuko Matsubara;Yasushi Sakurai
Kouki Kawabata;Yasuko Matsubara;Yasushi Sakurai
中科院分区:
其他
文献类型:
--
作者:
Kouki Kawabata;Yasuko Matsubara;Yasushi Sakurai

文献摘要

被引文献

相似文献

考虑到大量的多维数据流,例如物联网应用程序、金融和在线网络点击日志产生的数据流,我们如何发现典型模式并将其压缩成紧凑的模型?此外,我们如何在考虑从流媒体设置中发现的模式获得的信息的同时逐步区分多个模式?在本文中,我们提出了一个流算法,即StreamScope,旨在有效地发现直观的模式,随着时间的推移,从事件流演变。我们所提出的方法具有以下特性:(a)它是有效的:它对共同进化流的半无限集合进行操作,并将所有流总结成一组多个离散段,并按其相似性进行分组。(b)它是自动的:它可以自动和递增地识别这些模式,并在必要时为每个模式生成模型;(c)它是可扩展的:我们的方法的复杂性不取决于数据流的长度。我们在真实的数据流上进行的大量实验表明,StreamScope可以找到有意义的模式,并在计算时间和内存空间方面取得了很大的改进。
Given a large volume of multi-dimensional data streams, such as that produced by IoT applications, finance and online web-click logs, how can we discover typical patterns and compress them into compact models? In addition, how can we incrementally distinguish multiple patterns while considering the information obtained from a pattern found in a streaming setting? In this paper, we propose a streaming algorithm, namely StreamScope, that is designed to find intuitive patterns efficiently from event streams evolving over time. Our proposed method has the following properties: (a) it is effective: it operates on semi-infinite collections of co-evolving streams and summarizes all the streams into a set of multiple discrete segments grouped by their similarities. (b) it is automatic: it automatically and incrementally recognizes such patterns and generates models for each of them if necessary; (c) it is scalable: the complexity of our method does not depend on the length of the data streams. Our extensive experiments on real data streams demonstrate that StreamScope can find meaningful patterns and achieve great improvements in terms of computational time and memory space over its full batch method competitors.