Item Sets that Compress

Item Sets that Compress
复制标题

DOI:
10.1137/1.9781611972764.35
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
A. Siebes;Jilles Vreeken;M. Leeuwen
A. Siebes;Jilles Vreeken;M. Leeuwen
中科院分区:
其他
文献类型:
--
作者:
A. Siebes;Jilles Vreeken;M. Leeuwen

文献摘要

被引文献

相似文献

频繁项目集挖掘的主要问题之一是结果数量的爆炸性增长:很难找到最感兴趣的频繁项目集。这种爆炸的原因是大的频繁项目集描述了本质上相同的事务集。在本文中,我们使用MDL原理来处理这个问题:最好的频繁项目集集合是对数据库压缩最好的集合。针对这一问题,我们提出了四种启发式算法,实验表明,这些算法大大减少了频繁项目集的数量。此外,我们还展示了如何使用我们的方法来确定min-sup阈值的最佳值。
One of the major problems in frequent item set mining is the explosion of the number of results: it is difficult to find the most interesting frequent item sets. The cause of this explosion is that large sets of frequent item sets describe essentially the same set of transactions. In this paper we approach this problem using the MDL principle: the best set of frequent item sets is that set that compresses the database best. We introduce four heuristic algorithms for this task, and the experiments show that these algorithms give a dramatic reduction in the number of frequent item sets. Moreover, we show how our approach can be used to determine the best value for the min-sup threshold.