Stable Prediction With Leveraging Seed Variable

Stable Prediction With Leveraging Seed Variable
复制标题

DOI:
10.1109/tkde.2022.3169333
复制
发表时间:
2023-06
影响因子:
8.9
通讯作者:
Kun Kuang;Haotian Wang;Yue Liu;Ruoxuan Xiong;Runze Wu;Weiming Lu;Y. Zhuang;Fei Wu;Peng Cui;B. Li
Kun Kuang;Haotian Wang;Yue Liu;Ruoxuan Xiong;Runze Wu;Weiming Lu;Y. Zhuang;Fei Wu;Peng Cui;B. Li
中科院分区:
计算机科学2区
文献类型:
--
作者:
Kun Kuang;Haotian Wang;Yue Liu;Ruoxuan Xiong;Runze Wu;Weiming Lu;Y. Zhuang;Fei Wu;Peng Cui;B. Li

文献摘要

相似文献

In this paper, we focus on the problem of stable prediction across unknown test data, where the test distribution might be different from the training one and is always agnostic when model training. In such a case, previous machine learning methods might exploit subtly spurious correlations induced by non-causal variables in training data for prediction. Those spurious correlations can vary across datasets, leading to instability of prediction across unknown test data. To address this problem, we propose an algorithm based on conditional independence tests to screen out non-causal features and reduce spurious correlations by leveraging a seed variable. We show, both theoretically and with empirical experiments, that our algorithm can precisely screen out the isolated non-causal variables, which have no causal relationship with other variables, and remove the spurious correlations induced by them, increasing the stability of prediction across unknown test data. Extensive experiments on both synthetic and real-world datasets demonstrate that our algorithm outperforms state-of-the-art methods for stable prediction across unknown test data.