Balance-Subsampled Stable Prediction Across Unknown Test Data

Balance-Subsampled Stable Prediction Across Unknown Test Data
复制标题

DOI:
10.1145/3477052
复制
发表时间:
2021-10
期刊:
ACM Transactions on Knowledge Discovery from Data (TKDD)
影响因子:
--
通讯作者:
Kun Kuang;Hengtao Zhang;Runze Wu;Fei Wu;Y. Zhuang;Aijun Zhang
Kun Kuang;Hengtao Zhang;Runze Wu;Fei Wu;Y. Zhuang;Aijun Zhang
中科院分区:
其他
文献类型:
--
作者:
Kun Kuang;Hengtao Zhang;Runze Wu;Fei Wu;Y. Zhuang;Aijun Zhang

文献摘要

相似文献

在数据挖掘和机器学习中,通常假设训练和测试数据共享相同的总体分布。然而,由于样本选择偏差的存在,这一假设在实际应用中常常被违背,这可能导致从训练数据到测试数据的分布偏移。这种与模型无关的分布偏移通常会导致未知测试数据的预测不稳定。基于部分析因设计理论,提出了一种新的平衡下采样稳定预测(BSSP)算法。它将每个预测因子的明确效应与混杂变量隔离开来。设计理论分析表明,该方法能有效降低分布漂移引起的预测变量间的混杂效应,提高参数估计的准确性和对未知测试数据的预测稳定性.在合成数据集和真实数据集上的数值实验表明,我们的BSSP算法在未知测试数据上的稳定预测性能明显优于基线方法。
In data mining and machine learning, it is commonly assumed that training and test data share the same population distribution. However, this assumption is often violated in practice because of the sample selection bias, which might induce the distribution shift from training data to test data. Such a model-agnostic distribution shift usually leads to prediction instability across unknown test data. This article proposes a novel balance-subsampled stable prediction (BSSP) algorithm based on the theory of fractional factorial design. It isolates the clear effect of each predictor from the confounding variables. A design-theoretic analysis shows that the proposed method can reduce the confounding effects among predictors induced by the distribution shift, improving both the accuracy of parameter estimation and the stability of prediction across unknown test data. Numerical experiments on synthetic and real-world datasets demonstrate that our BSSP algorithm can significantly outperform the baseline methods for stable prediction across unknown test data.