Scalable techniques for mining causal structures

Scalable techniques for mining causal structures
复制标题

DOI:
10.1023/a:1009891813863
复制
发表时间:
2000-07-01
影响因子:
4.8
通讯作者:
Ullman, J
Ullman, J
中科院分区:
计算机科学3区
文献类型:
--
作者:
Silverstein, C;Brin, S;Ullman, J

文献摘要

被引文献

相似文献

在购物篮数据中挖掘关联规则是一个富有成果的研究领域。诸如条件概率(置信度)和相关性之类的度量已被用于推断“项目A的存在意味着项目B的存在”形式的规则。“然而,这样的规则仅表明A和B之间的统计关系。它们并没有具体说明这种关系的性质:是A的存在导致B的存在,还是匡威,或者是其他一些属性或现象导致两者同时出现。在应用中,了解这种因果关系对于加强理解和实现变革非常有用。虽然区分因果关系和相关性是一个真正困难的问题,但最近在统计学和贝叶斯学习方面的工作提供了一些攻击途径。在这些领域中,目标一般是学习完整的因果模型,这基本上是不可能学习的大规模数据挖掘应用程序中有大量的variables.In本文中,我们考虑的问题,确定因果关系,而不仅仅是协会,当挖掘市场篮子数据。我们确定了一些问题与贝叶斯学习思想挖掘大型数据库的直接应用,关于算法的可扩展性和适当的统计技术,并介绍了一些初步的想法来处理这些问题。我们目前的实验结果,从应用我们的算法在几个大的,真实世界的数据集。结果表明,这里提出的方法是计算上可行的,并成功地识别有趣的因果结构。一个有趣的结果是,推断因果关系的缺乏可能比推断因果关系更容易,这些信息有助于防止错误的决策。
Mining for association rules in market basket data has proved a fruitful area of research. Measures such as conditional probability (confidence) and correlation have been used to infer rules of the form "the existence of item A implies the existence of item B." However, such rules indicate only a statistical relationship between A and B. They do not specify the nature of the relationship: whether the presence of A causes the presence of B, or the converse, or some other attribute or phenomenon causes both to appear together. In applications, knowing such causal relationships is extremely useful for enhancing understanding and effecting change. While distinguishing causality from correlation is a truly difficult problem, recent work in statistics and Bayesian learning provide some avenues of attack. In these fields, the goal has generally been to learn complete causal models, which are essentially impossible to learn in large-scale data mining applications with a large number of variables.In this paper, we consider the problem of determining casual relationships, instead of mere associations, when mining market basket data. We identify some problems with the direct application of Bayesian learning ideas to mining large databases, concerning both the scalability of algorithms and the appropriateness of the statistical techniques, and introduce some initial ideas for dealing with these problems. We present experimental results from applying our algorithms on several large, real-world data sets. The results indicate that the approach proposed here is both computationally feasible and successful in identifying interesting causal structures. An interesting outcome is that it is perhaps easier to infer the lack of causality than to infer causality, information that is useful in preventing erroneous decision making.