Scalable and Hybrid Ensemble-Based Causality Discovery

Scalable and Hybrid Ensemble-Based Causality Discovery
复制标题

DOI:
10.1109/smds49396.2020.00016
复制
发表时间:
2020-10
期刊:
2020 IEEE International Conference on Smart Data Services (SMDS)
影响因子:
--
通讯作者:
Pei Guo;Achuna Ofonedu;Jianwu Wang
Pei Guo;Achuna Ofonedu;Jianwu Wang
中科院分区:
其他
文献类型:
--
作者:
Pei Guo;Achuna Ofonedu;Jianwu Wang

文献摘要

被引文献

相似文献

因果关系发现挖掘系统中不同变量之间的因果关系,已广泛应用于许多学科,包括气候学和神经科学。为了发现因果关系,人们提出了许多数据驱动的因果关系发现方法,如格兰杰因果关系、PCMCI和动态贝叶斯网络。许多因果关系发现方法挖掘时间序列数据并生成有向因果图,其中每个图边表示两个连接图节点之间的因果关系。我们对不同因果关系发现方法与现实世界气候数据的基准测试表明,由于内部学习机制的差异,这些方法通常会在相同的输入数据集上产生完全不同的因果关系结果。同时,几乎每个学科的可用数据都在不断增加,这使得使用现有的因果关系发现算法在合理的时间内产生因果关系结果变得越来越困难。为了解决这两个问题,本文利用数据划分和集成技术,提出了一个两阶段混合因果关系集成框架。框架首先对分区数据进行第一阶段的数据集成,然后根据数据集成结果进行第二阶段的算法集成。为了实现可扩展性,我们通过Spark大数据分析引擎进一步并行化集成方法。实验表明,本文提出的方法在分布式计算环境下通过集成获得了良好的精度,并通过数据并行化获得了较高的可扩展性。
Causality discovery mines cause-effect relationships among different variables of a system and has been widely used in many disciplines including climatology and neuroscience. To discover causal relationships, many data-driven causality discovery methods, e.g., Granger causality, PCMCI and Dynamic Bayesian Network, have been proposed. Many of these causality discovery approaches mine time series data and generate a directed causality graph where each graph edge denotes a cause-effect relationship between the two connected graph nodes. Our benchmarking of different causality discovery approaches with real-world climate data shows these approaches often generate quite different causality results with the same input dataset due to their internal learning mechanism differences. Meanwhile, there are ever-increasing available data in virtually every discipline, which makes it more and more difficult to use existing causality discovery algorithms to produce causality results within reasonable time. To address these two challenges, this paper utilizes data partitioning and ensemble techniques, and proposes a two-phase hybrid causality ensemble framework. The framework first conducts phase 1 data ensemble for partitioned data and then conducts phase 2 algorithm ensemble from data ensemble results. To achieve scalability, we further parallelize the ensemble approaches via the Spark big data analytics engine. Our experiments show that our proposed approaches achieve good accuracy through ensemble and high scalability through data-parallelization in distributed computing environments.