Out-of-core frequent pattern mining on a commodity PC

Out-of-core frequent pattern mining on a commodity PC
复制标题

DOI:
10.1145/1150402.1150416
复制
发表时间:
2006-08
期刊:
--
影响因子:
--
通讯作者:
G. Buehrer;S. Parthasarathy;A. Ghoting
G. Buehrer;S. Parthasarathy;A. Ghoting
中科院分区:
其他
文献类型:
--
作者:
G. Buehrer;S. Parthasarathy;A. Ghoting

文献摘要

被引文献

相似文献

在这项工作中,我们集中在大型核外数据集上的频繁项集挖掘问题。在介绍了现有的核外频繁项集挖掘算法的特点及其不足之后,我们介绍了我们的高效、高度可扩展的解决方案。我们的技术是在FPGrowth算法的背景下提出的,它涉及几种新颖的I/O敏感优化,如基于近似哈希的排序和分块,并利用了商用计算机中最新的体系结构改进,如位处理。我们在非常大的数据集(高达75 GB)上对建议的优化进行了评估,结果显示它们的执行时间缩短了400倍以上。最后,我们讨论了这项研究在其他模式挖掘挑战的背景下的影响,如序列挖掘和图挖掘。
In this work we focus on the problem of frequent itemset mining on large, out-of-core data sets. After presenting a characterization of existing out-of-core frequent itemset mining algorithms and their drawbacks, we introduce our efficient, highly scalable solution. Presented in the context of the FPGrowth algorithm, our technique involves several novel I/O-conscious optimizations, such as approximate hash-based sorting and blocking, and leverages recent architectural advancements in commodity computers, such as 64-bit processing. We evaluate the proposed optimizations on truly large data sets,up to 75GB, and show they yield greater than a 400-fold execution time improvement. Finally, we discuss the impact of this research in the context of other pattern mining challenges, such as sequence mining and graph mining.