Efficient mining frequent itemsets algorithms

Efficient mining frequent itemsets algorithms
复制标题

DOI:
10.1007/s13042-013-0172-6
复制
发表时间:
2013-05
影响因子:
5.6
通讯作者:
M. Mohamed;Mohammed M. Darwieesh
M. Mohamed;Mohammed M. Darwieesh
中科院分区:
计算机科学3区
文献类型:
--
作者:
M. Mohamed;Mohammed M. Darwieesh

文献摘要

被引文献

相似文献

高效的频繁项集挖掘算法对于关联规则挖掘以及其他数据挖掘任务都是至关重要的。众所周知,countTable是利用子集属性将事务数据库压缩为新的更低的出现项表示的最重要的工具之一。该技术的最大问题之一是候选生成和测试处理的成本,这是发现关联规则的两个最重要的步骤。在本文中,我们已经开发了这种方法,以避免昂贵的候选生成和测试处理完全。此外,所提出的方法还压缩了关键信息的所有项目集,最大长度的频繁项目集,最小长度的频繁项目集,避免昂贵的,重复的数据库扫描。本文提出的算法名为CountTableFI和BinaryCountTableF,该算法与Apriori算法以及所有由Apriori算法扩展而来的算法有显著的区别。该算法的思想是在事务的表示中,我们用二进制数和十进制数表示所有事务,因此使用子集和恒等集属性是简单和快速的。一个全面的性能研究表明,我们的技术是有效的,可扩展性与其他方法相比。
Efficient algorithms for mining frequent itemsets are crucial for mining association rules as well as for many other data mining tasks. It is well known that countTable is one of the most important facility to employ subsets property for compressing the transaction database to new lower representation of occurrences items. One of the biggest problem in this technique is the cost of candidate generation and test processing which are the two most important steps to find association rules. In this paper, we have developed this method to avoid the costly candidate-generation-and-test processing completely. Moreover, the proposed methods also compress crucial information about all itemsets, maximal length frequent itemsets, minimal length frequent itemsets, avoid expensive, and repeated database scans. The proposed named CountTableFI and BinaryCountTableF are presented, the algorithm has significant difference from the Apriori and all other algorithms extended from Apriori. The idea behind this algorithm is in the representation of the transactions, where, we represent all transactions in binary number and decimal number, so it is simple and fast to use subset and identical set properties. A comprehensive performance study shows that our techniques are efficient and scalable comparing with other methods.