Stratified Learning: a general-purpose statistical method for improved learning under Covariate Shift

Stratified Learning: a general-purpose statistical method for improved learning under Covariate Shift
复制标题

DOI:
10.1002/sam.11643
复制
发表时间:
2021-06
期刊:
Statistical Analysis and Data Mining: The ASA Data Science Journal
影响因子:
--
通讯作者:
Maximilian Autenrieth;D. V. Dyk;R. Trotta;D. Stenning
Maximilian Autenrieth;D. V. Dyk;R. Trotta;D. Stenning
中科院分区:
其他
文献类型:
--
作者:
Maximilian Autenrieth;D. V. Dyk;R. Trotta;D. Stenning

文献摘要

相似文献

我们提出了一种简单的,统计学原则和理论上合理的方法来改进监督学习,当训练集不具有代表性时,这种情况称为协变量移位。我们建立在因果推断的一个完善的方法,并表明协变量变化的影响可以通过对倾向评分的调节来减少或消除。在实践中,这是通过将学习者拟合在通过基于估计的倾向分数划分数据而构建的层中来实现的,从而导致近似平衡的协变量和大大改进的目标预测。我们将这种方法称为分层学习或StratLearn。我们证明了这种通用方法对宇宙学中两个当代研究问题的有效性,优于最先进的重要性加权方法。我们在更新的“超新星光度分类挑战”中获得了最佳报告的AUC(0.958),并且我们改进了斯隆数字巡天(SDSS)数据中星系红移的现有条件密度估计。
We propose a simple, statistically principled, and theoretically justified method to improve supervised learning when the training set is not representative, a situation known as covariate shift. We build upon a well‐established methodology in causal inference and show that the effects of covariate shift can be reduced or eliminated by conditioning on propensity scores. In practice, this is achieved by fitting learners within strata constructed by partitioning the data based on the estimated propensity scores, leading to approximately balanced covariates and much‐improved target prediction. We refer to the overall method as Stratified Learning, or StratLearn. We demonstrate the effectiveness of this general‐purpose method on two contemporary research questions in cosmology, outperforming state‐of‐the‐art importance weighting methods. We obtain the best‐reported AUC (0.958) on the updated “Supernovae photometric classification challenge,” and we improve upon existing conditional density estimation of galaxy redshift from Sloan Digital Sky Survey (SDSS) data.