PAC Prediction Sets Under Covariate Shift

PAC Prediction Sets Under Covariate Shift
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Sangdon Park;Edgar Dobriban;Insup Lee;O. Bastani
Sangdon Park;Edgar Dobriban;Insup Lee;O. Bastani
中科院分区:
其他
文献类型:
--
作者:
Sangdon Park;Edgar Dobriban;Insup Lee;O. Bastani

文献摘要

相似文献

现代机器学习面临的一个重要挑战是如何严格量化模型预测的不确定性。当可能使预测模型无效的基础数据分布发生变化时,传达不确定性尤为重要。然而,在存在此类转移的情况下,大多数现有的不确定性定量算法分解。我们提出了一种新颖的方法,该方法通过构建\ emph {可能是正确的(PAC)}的预测集在存在协变量的情况下来解决这一挑战。我们的方法着重于从源分布(我们标记为培训示例)到目标分布(我们想要量化不确定性)的协变量转移的设置。我们的算法假设给定重要的权重编码在协变量转移下训练示例的概率如何变化。实际上,重要的权重通常需要估计;因此,我们将算法扩展到给予重要性权重的置信区间的设置。我们证明了基于域内和成像网的协变量转移的有效性。我们的算法满足PAC的约束,并在始终满足PAC约束的方法中给出的预测集具有最小的平均归一化大小。
An important challenge facing modern machine learning is how to rigorously quantify the uncertainty of model predictions. Conveying uncertainty is especially important when there are changes to the underlying data distribution that might invalidate the predictive model. Yet, most existing uncertainty quantification algorithms break down in the presence of such shifts. We propose a novel approach that addresses this challenge by constructing \emph{probably approximately correct (PAC)} prediction sets in the presence of covariate shift. Our approach focuses on the setting where there is a covariate shift from the source distribution (where we have labeled training examples) to the target distribution (for which we want to quantify uncertainty). Our algorithm assumes given importance weights that encode how the probabilities of the training examples change under the covariate shift. In practice, importance weights typically need to be estimated; thus, we extend our algorithm to the setting where we are given confidence intervals for the importance weights. We demonstrate the effectiveness of our approach on covariate shifts based on DomainNet and ImageNet. Our algorithm satisfies the PAC constraint, and gives prediction sets with the smallest average normalized size among approaches that always satisfy the PAC constraint.