Fast vertical mining using diffsets

Fast vertical mining using diffsets
复制标题

DOI:
10.1145/956750.956788
复制
发表时间:
2003-08
期刊:
--
影响因子:
--
通讯作者:
Mohammed J. Zaki;K. Gouda
Mohammed J. Zaki;K. Gouda
中科院分区:
其他
文献类型:
--
作者:
Mohammed J. Zaki;K. Gouda

文献摘要

被引文献

相似文献

最近已经提出了一些垂直挖掘算法来进行关联挖掘,这些算法被证明是非常有效的,并且通常优于水平挖掘方法。垂直格式的主要优点是支持通过对事务ID(TID)的交集操作进行快速频率计数和自动修剪不相关的数据。本文提出了一种新的垂直数据表示方法Diffset,该方法只跟踪候选模式的TID之间的差异,避免了频繁模式的产生。我们表明,差值大大减少了存储中间结果所需的内存大小。我们展示了当差异被结合到以前的垂直挖掘方法中时,如何显著提高性能。
A number of vertical mining algorithms have been proposed recently for association mining, which have shown to be very effective and usually outperform horizontal approaches. The main advantage of the vertical format is support for fast frequency counting via intersection operations on transaction ids (tids) and automatic pruning of irrelevant data. The main problem with these approaches is when intermediate results of vertical tid lists become too large for memory, thus affecting the algorithm scalability.In this paper we present a novel vertical data representation called Diffset, that only keeps track of differences in the tids of a candidate pattern from its generating frequent patterns. We show that diffsets drastically cut down the size of memory required to store intermediate results. We show how diffsets, when incorporated into previous vertical mining methods, increase the performance significantly.