Exploratory Causal Analysis of Open Data: Explanation Generation and Confounder Identification

Exploratory Causal Analysis of Open Data: Explanation Generation and Confounder Identification
复制标题

开放数据的探索性因果分析:解释生成和混杂因素识别

DOI:
10.20965/jaciii.2020.p0142
复制
发表时间:
2020
期刊:
J. Adv. Comput. Intell. Intell. Informatics
影响因子:
--
通讯作者:
M. Kurihara
M. Kurihara
中科院分区:
--
文献类型:
--
作者:
Jing Song;S. Oyama;M. Kurihara

文献摘要

参考文献

被引文献

相似文献

开放数据在各个领域变得越来越可用,许多组织依赖于根据数据做出决策。这样的决策需要小心区分相关性和因果关系。在数据分析任务中,由于未观察到的混杂因素,因果关系分析特别复杂。例如,为了正确分析两个变量之间的因果关系,应该考虑第三个变量可能的混杂效应。然而,在开放数据环境中,很难提前考虑所有可能的混杂因素。在本文中,我们提出了一个开放数据的探索性因果分析的框架,其中可能的混杂变量收集和增量测试从大量的开放数据。据作者所知,尚未提出在因果分析过程中纳入可能混杂因素数据的框架。本文展示了一种扩展因果结构并生成合理因果关系的独创性方法。拟议的框架占因果分析中可能的混淆的影响,首先使用众包平台收集变量之间的相关性的解释。然后使用自然语言处理方法提取关键词。该框架根据提取的关键字搜索相关的开放数据。最后,使用几种自动因果分析方法对收集到的解释进行测试。我们使用世界银行和日本政府的公开数据进行了实验。实验结果证实,所提出的框架,使因果分析,同时考虑可能的混杂因素的影响。
Open data are becoming increasingly available in various domains, and many organizations rely on making decisions according to data. Such decision making requires care to distinguish between correlations and causal relationships. Among data analysis tasks, causal relationship analysis is especially complex because of unobserved confounders. For example, to correctly analyze the causal relationship between two variables, the possible confounding effect of a third variable should be considered. In the open-data environment, however, it is difficult to consider all possible confounders in advance. In this paper, we propose a framework for exploratory causal analysis of open data, in which possible confounding variables are collected and incrementally tested from a large volume of open data. To the extent of the authors’ knowledge, no framework has been proposed to incorporate data for possible confounders in causal analysis process. This paper shows an original way to expand causal structures and generate reasonable causal relationships. The proposed framework accounts for the effect of possible confounding in causal analysis by first using a crowdsourcing platform to collect explanations of the correlation between variables. Keywords are then extracted using natural language processing methods. The framework searches the related open data according to the extracted keywords. Finally, the collected explanations are tested using several automated causal analysis methods. We conducted experiments using open data from the World Bank and the Japanese government. The experimental results confirmed that the proposed framework enables causal analysis while considering the effects of possible confounders.
DOI: --
发表时间: 2014
期刊:
影响因子: --
作者:
Lei Duan;Satoshi Oyama;Haruhiko Sato;and Masahito Kurihara
通讯作者: and Masahito Kurihara
DOI: 10.1016/j.ijar.2008.02.006
发表时间: 2008-10-01
影响因子: 3.9
作者:
Hoyer, Patrik O.;Shimizu, Shohei;Palviainen, Markus
通讯作者: Palviainen, Markus