Boosting association rule mining in large datasets via Gibbs sampling

Boosting association rule mining in large datasets via Gibbs sampling
复制标题

DOI:
10.1073/pnas.1604553113
复制
发表时间:
2016-05-03
影响因子:
11.1
通讯作者:
Wu, Yuehua
Wu, Yuehua
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Qian, Guoqi;Rao, Calyampudi Radhakrishna;Wu, Yuehua

文献摘要

被引文献

相似文献

当前从事务数据中挖掘关联规则的算法大多是确定性的和枚举的。如果不采取限制搜索空间的措施,即使对只包含几百个事务项的数据集进行挖掘,它们在计算上也是难以处理的。在本文中,我们开发了一种Gibbs-sampling-induced随机搜索程序,从项目集空间中随机抽取关联规则,并从样本生成的约简事务数据集中进行规则挖掘。同时,提出了一种通用规则重要性度量来指导随机搜索,使得随机生成的关联规则构成一个遍历马尔可夫链,从而可以在极限概率为1的情况下从简化的数据集中发现项目集空间中总体上最重要的规则。在模拟研究和一个真实的基因组数据示例中,我们展示了如何通过集成使用随机搜索和Apriori算法来增强关联规则挖掘。
Current algorithms for association rule mining from transaction data are mostly deterministic and enumerative. They can be computationally intractable even for mining a dataset containing just a few hundred transaction items, if no action is taken to constrain the search space. In this paper, we develop a Gibbs-sampling-induced stochastic search procedure to randomly sample association rules from the itemset space, and perform rule mining from the reduced transaction dataset generated by the sample. Also a general rule importance measure is proposed to direct the stochastic search so that, as a result of the randomly generated association rules constituting an ergodic Markov chain, the overall most important rules in the itemset space can be uncovered from the reduced dataset with probability 1 in the limit. In the simulation study and a real genomic data example, we show how to boost association rule mining by an integrated use of the stochastic search and the Apriori algorithm.