Boosting association rule mining in large datasets via Gibbs sampling
Boosting association rule mining in large datasets via Gibbs sampling
复制标题
DOI:
10.1073/pnas.1604553113
复制
发表时间:
2016-05-03
影响因子:
11.1
通讯作者:
Wu, Yuehua
中科院分区:
文献类型:
--
作者:
Qian, Guoqi;Rao, Calyampudi Radhakrishna;Wu, Yuehua
Current algorithms for association rule mining from transaction data are mostly deterministic and enumerative. They can be computationally intractable even for mining a dataset containing just a few hundred transaction items, if no action is taken to constrain the search space. In this paper, we develop a Gibbs-sampling-induced stochastic search procedure to randomly sample association rules from the itemset space, and perform rule mining from the reduced transaction dataset generated by the sample. Also a general rule importance measure is proposed to direct the stochastic search so that, as a result of the randomly generated association rules constituting an ergodic Markov chain, the overall most important rules in the itemset space can be uncovered from the reduced dataset with probability 1 in the limit. In the simulation study and a real genomic data example, we show how to boost association rule mining by an integrated use of the stochastic search and the Apriori algorithm.