Collaborative causal inference on distributed data

Collaborative causal inference on distributed data
复制标题

DOI:
10.1016/j.eswa.2023.123024
复制
发表时间:
2022-08
期刊:
Expert Syst. Appl.
影响因子:
--
通讯作者:
Y. Kawamata;Ryoki Motai;Yukihiko Okada;A. Imakura;T. Sakurai
Y. Kawamata;Ryoki Motai;Yukihiko Okada;A. Imakura;T. Sakurai
中科院分区:
其他
文献类型:
--
作者:
Y. Kawamata;Ryoki Motai;Yukihiko Okada;A. Imakura;T. Sakurai

文献摘要

相似文献

近年来,分布式数据隐私保护的因果推理技术的发展得到了相当大的关注。现有的许多分布式数据的方法都集中在解决受试者(样本)的缺乏,只能减少估计治疗效果的随机误差。在这项研究中,我们提出了一个数据协作准实验(DC-QE),解决了缺乏主题和协变量,减少随机误差和估计偏差。我们的方法包括从本地各方的私人数据中构建降维的中间表示,共享中间表示而不是隐私保护的私人数据,从共享的中间表示中估计倾向分数,最后,从倾向分数估计治疗效果。通过对人工和真实世界数据的数值实验,我们证实了我们的方法比单独分析的估计结果更好。虽然降维在私有数据中丢失了一些信息并导致性能下降,但我们观察到,与多方共享中间表示以解决主题和协变量的缺乏足以提高性能,以克服降维引起的性能下降。虽然外部有效性不一定得到保证,我们的研究结果表明,DC-QE是一个很有前途的方法。随着我们的方法的广泛使用,中间表示可以作为开放数据发布,以帮助研究人员找到因果关系并积累知识库。
In recent years, the development of technologies for causal inference with privacy preservation of distributed data has gained considerable attention. Many existing methods for distributed data focus on resolving the lack of subjects (samples) and can only reduce random errors in estimating treatment effects. In this study, we propose a data collaboration quasi-experiment (DC-QE) that resolves the lack of both subjects and covariates, reducing random errors and biases in the estimation. Our method involves constructing dimensionality-reduced intermediate representations from private data from local parties, sharing intermediate representations instead of private data for privacy preservation, estimating propensity scores from the shared intermediate representations, and finally, estimating the treatment effects from propensity scores. Through numerical experiments on both artificial and real-world data, we confirm that our method leads to better estimation results than individual analyses. While dimensionality reduction loses some information in the private data and causes performance degradation, we observe that sharing intermediate representations with many parties to resolve the lack of subjects and covariates sufficiently improves performance to overcome the degradation caused by dimensionality reduction. Although external validity is not necessarily guaranteed, our results suggest that DC-QE is a promising method. With the widespread use of our method, intermediate representations can be published as open data to help researchers find causalities and accumulate a knowledge base.