Efficient mining method for retrieving sequential patterns over online data streams

Efficient mining method for retrieving sequential patterns over online data streams
复制标题

DOI:
10.1177/0165551505055405
复制
发表时间:
2005-01-01
影响因子:
2.4
通讯作者:
Lee, WS
Lee, WS
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chang, JH;Lee, WS

文献摘要

被引文献

相似文献

随着数据挖掘在信息科学各个领域的应用,在以往的研究中提出了各种挖掘方法。最近,在这些领域中,数据采用连续数据流的形式,而不是有限存储的数据集。本文提出了一种在线序列数据流上序列模式的挖掘方法,该方法可用于检索数据流中的嵌入式知识。该方法可以在允许错误的情况下,最大限度地减少挖掘过程的内存使用,并支持在内存使用和挖掘精度之间进行灵活的权衡。然而,通过对序列计数的精确估计方法,该方法考虑了项目的排序信息,使误差最小化。该方法可以在短时间内捕获序列数据流中最近的变化,通过一种衰变机制优雅地丢弃可能不再有用的旧信息。
With the usefulness of data mining in various fields of information science, various mining methods have been proposed in previous research. Recently, in these fields, data has taken the form of continuous data streams rather than finite stored data sets. In this paper, a mining method of sequential patterns over an online sequence data stream is proposed, which is useful for retrieving embedded knowledge in the data stream. The proposed method can minimize memory usage of the mining process while an error is allowed in its mining result, and supports flexible trade-off between memory usage and mining accuracy. However, the error is minimized by an accurate estimation method for the count of a sequence, which considers the ordering information of items. The proposed method can catch a recent change in a sequence data stream in a short time, by a decaying mechanism gracefully discarding old information that may be no longer useful.