Finding interesting associations without support pruning

Finding interesting associations without support pruning
复制标题

DOI:
10.1109/icde.2000.839448
复制
发表时间:
2000-02
期刊:
Proceedings of 16th International Conference on Data Engineering (Cat. No.00CB37073)
影响因子:
--
通讯作者:
E. Cohen;Mayur Datar;S. Fujiwara;A. Gionis;P. Indyk;R. Motwani;J. Ullman;Cheng Yang
E. Cohen;Mayur Datar;S. Fujiwara;A. Gionis;P. Indyk;R. Motwani;J. Ullman;Cheng Yang
中科院分区:
其他
文献类型:
--
作者:
E. Cohen;Mayur Datar;S. Fujiwara;A. Gionis;P. Indyk;R. Motwani;J. Ullman;Cheng Yang

文献摘要

被引文献

相似文献

迄今为止,关联规则挖掘一直依赖于高支持度的条件来有效地完成其工作。特别地,众所周知的先验算法仅在感兴趣的唯一规则是非常频繁出现的关系时才有效。然而,有一些应用程序,如数据挖掘,识别相似的Web文档,聚类和协同过滤,其中感兴趣的规则在数据中的实例相对较少。在这些情况下,我们必须寻找高度相关的项目,甚至可能是罕见项目之间的因果关系。我们开发了一个家庭的算法来解决这个问题,采用随机抽样和哈希技术的组合。我们提供了一个分析的算法开发和进行实验,真实的和合成数据,以获得比较性能分析。
Association rule mining has heretofore relied on the condition of high support to do its work efficiently. In particular, the well-known a-priori algorithm is only effective when the only rules of interest are relationships that occur very frequently. However, there are a number of applications, such as data mining, identification of similar Web documents, clustering and collaborative filtering, where the rules of interest have comparatively few instances in the data. In these cases, we must look for highly correlated items, or possibly even causal relationships between infrequent items. We develop a family of algorithms for solving this problem, employing a combination of random sampling and hashing techniques. We provide an analysis of the algorithms developed and conduct experiments on real and synthetic data to obtain a comparative performance analysis.