Frequent pattern mining with uncertain data

Frequent pattern mining with uncertain data
复制标题

DOI:
10.1145/1557019.1557030
复制
发表时间:
2009-06
影响因子:
2.8
通讯作者:
C. Aggarwal;Yan Li;Jianyong Wang;Jing Wang
C. Aggarwal;Yan Li;Jianyong Wang;Jing Wang
中科院分区:
工程技术4区
文献类型:
--
作者:
C. Aggarwal;Yan Li;Jianyong Wang;Jing Wang

文献摘要

被引文献

相似文献

研究了不确定数据下的频繁模式挖掘问题。我们将展示如何将广泛的算法类扩展到不确定数据集。特别是,我们将研究候选的生成和测试算法,超结构算法和基于模式增长的算法。我们的一个有见地的观察是,在不确定的情况下,与确定的情况相比,不同类别的算法的实验行为是非常不同的。特别是,超结构和候选生成和测试算法比基于树的算法表现得更好。这种反直觉的行为是从算法设计的角度观察问题的不确定性变化的一个重要现象。我们将在许多真实和合成数据集上测试该方法,并展示我们的两种方法相对于竞争技术的有效性。
This paper studies the problem of frequent pattern mining with uncertain data. We will show how broad classes of algorithms can be extended to the uncertain data setting. In particular, we will study candidate generate-and-test algorithms, hyper-structure algorithms and pattern growth based algorithms. One of our insightful observations is that the experimental behavior of different classes of algorithms is very different in the uncertain case as compared to the deterministic case. In particular, the hyper-structure and the candidate generate-and-test algorithms perform much better than tree-based algorithms. This counter-intuitive behavior is an important observation from the perspective of algorithm design of the uncertain variation of the problem. We will test the approach on a number of real and synthetic data sets, and show the effectiveness of two of our approaches over competitive techniques.