Causal Datasheet for Datasets: An Evaluation Guide for Real-World Data Analysis and Data Collection Design Using Bayesian Networks.

Causal Datasheet for Datasets: An Evaluation Guide for Real-World Data Analysis and Data Collection Design Using Bayesian Networks.
复制标题

DOI:
10.3389/frai.2021.612551
复制
发表时间:
2021
影响因子:
4
通讯作者:
Quadrianto N
Quadrianto N
中科院分区:
其他
文献类型:
--
作者:
Butcher B;Huang VS;Robinson C;Reffin J;Sgaier SK;Charles G;Quadrianto N

文献摘要

参考文献

被引文献

相似文献

开发解决现实问题的数据驱动解决方案需要了解这些问题的原因以及它们的相互作用如何影响结果-通常只有观察数据。因果贝叶斯网络(BN)是一种从观测数据中发现因果关系并将其表示为有向无环图(DAG)的有效方法。BNs可能对中低收入国家的全球健康研究特别有用,这些国家有越来越多的观测数据可用于政策制定,项目评估和干预设计。然而,BN尚未被全球卫生专业人员广泛采用,在现实世界的应用中,对BN结果的信心通常仍然不足。这部分是由于无法验证一些基本事实,因为真正的DAG不可用。如果学习的DAG与预先存在的域原则相冲突,则这尤其成问题。在这里,我们概念化并展示了“因果数据表”的想法,该数据表可以近似并记录给定数据集的BN性能预期,旨在为从业者提供信心和样本量要求。为了生成这种因果数据表的结果,开发了一种工具,该工具可以生成合成贝叶斯网络及其相关的合成数据集,以模拟真实世界的数据集。著名的结构学习算法和一个新的实现的OrderMCMC方法使用商归一化最大似然分数给出的结果进行了记录。这些结果用于填充因果数据表,并根据预期性能是否满足用户定义的阈值提出建议。我们介绍了我们在创建因果数据表方面的经验,以帮助在研究过程的不同阶段做出分析决策。首先,部署了一个,以帮助确定适当的样本量的计划在印度中央邦的性健康和生殖健康的研究。第二,我们创建了一个评估表,以评估我们在印度北方邦进行的一项现有孕产妇健康调查的绩效。第三,我们验证了生成的性能估计,并调查了众所周知的ALARM数据集的当前限制。我们的经验证明了因果数据表的实用性,它可以帮助全球卫生从业者在应用BN时获得更多的信心。
Developing data-driven solutions that address real-world problems requires understanding of these problems’ causes and how their interaction affects the outcome–often with only observational data. Causal Bayesian Networks (BN) have been proposed as a powerful method for discovering and representing the causal relationships from observational data as a Directed Acyclic Graph (DAG). BNs could be especially useful for research in global health in Lower and Middle Income Countries, where there is an increasing abundance of observational data that could be harnessed for policy making, program evaluation, and intervention design. However, BNs have not been widely adopted by global health professionals, and in real-world applications, confidence in the results of BNs generally remains inadequate. This is partially due to the inability to validate against some ground truth, as the true DAG is not available. This is especially problematic if a learned DAG conflicts with pre-existing domain doctrine. Here we conceptualize and demonstrate an idea of a “Causal Datasheet” that could approximate and document BN performance expectations for a given dataset, aiming to provide confidence and sample size requirements to practitioners. To generate results for such a Causal Datasheet, a tool was developed which can generate synthetic Bayesian networks and their associated synthetic datasets to mimic real-world datasets. The results given by well-known structure learning algorithms and a novel implementation of the OrderMCMC method using the Quotient Normalized Maximum Likelihood score were recorded. These results were used to populate the Causal Datasheet, and recommendations could be made dependent on whether expected performance met user-defined thresholds. We present our experience in the creation of Causal Datasheets to aid analysis decisions at different stages of the research process. First, one was deployed to help determine the appropriate sample size of a planned study of sexual and reproductive health in Madhya Pradesh, India. Second, a datasheet was created to estimate the performance of an existing maternal health survey we conducted in Uttar Pradesh, India. Third, we validated generated performance estimates and investigated current limitations on the well-known ALARM dataset. Our experience demonstrates the utility of the Causal Datasheet, which can help global health practitioners gain more confidence when applying BNs.
DOI: 10.1136/bmjgh-2020-002340
发表时间: 2020-10
期刊: BMJ global health
影响因子: 8.1
作者:
Huang VS;Morris K;Jain M;Ramesh BM;Kemp H;Blanchard J;Isac S;Sarkar B;Gothalwal V;Namasivayam V;Kumar P;Sgaier SK
通讯作者: Sgaier SK