RLSbench: Domain Adaptation Under Relaxed Label Shift

RLSbench: Domain Adaptation Under Relaxed Label Shift
复制标题

DOI:
10.48550/arxiv.2302.03020
复制
发表时间:
2023-02
期刊:
--
影响因子:
--
通讯作者:
S. Garg;Nick Erickson;J. Sharpnack;Alexander J. Smola;Sivaraman Balakrishnan;Zachary Chase Lipton
S. Garg;Nick Erickson;J. Sharpnack;Alexander J. Smola;Sivaraman Balakrishnan;Zachary Chase Lipton
中科院分区:
其他
文献类型:
--
作者:
S. Garg;Nick Erickson;J. Sharpnack;Alexander J. Smola;Sivaraman Balakrishnan;Zachary Chase Lipton

文献摘要

被引文献

相似文献

尽管出现了标签转移下的领域适应原则方法,但它们对类别条件分布转移的敏感性尚不明确。与此同时,常用的深度领域自适应启发式算法在标签比例发生变化时往往会出现问题。虽然有几篇论文修改了这些启发式方法,试图处理标签比例的变化,但评估标准、数据集和基线的不一致性使得衡量当前的最佳实践变得困难。在本文中,我们介绍了RLSbench,这是一个大规模的宽松标签移动基准,由$>$500分布移动对组成,跨越视觉、表格和语言模式,具有不同的标签比例。现有的基准主要关注类别条件$p(x|y)$的变化,与之不同,我们的基准还关注标签边际变化。首先,我们评估了13种流行的领域自适应方法,表明在标签比例变化下,失败的范围比以前所知的要广泛。接下来,我们开发了一种有效的两步元算法,该算法与大多数领域自适应启发式算法兼容:(i)在每个epoch对数据进行伪平衡;(ii)根据目标标签分布估计调整最终分类器。元算法改进了现有的领域自适应启发式方法,在大标签比例变化下,通常提高2- 10%的精度点,而在标签比例不变化时,效果最小($<$0.5\%)。我们希望这些发现和RLSbench的可用性将鼓励研究人员在宽松的标签转换设置中严格评估所提出的方法。代码可在https://github.com/acmi-lab/RLSbench上公开获取。
Despite the emergence of principled methods for domain adaptation under label shift, their sensitivity to shifts in class conditional distributions is precariously under explored. Meanwhile, popular deep domain adaptation heuristics tend to falter when faced with label proportions shifts. While several papers modify these heuristics in attempts to handle label proportions shifts, inconsistencies in evaluation standards, datasets, and baselines make it difficult to gauge the current best practices. In this paper, we introduce RLSbench, a large-scale benchmark for relaxed label shift, consisting of $>$500 distribution shift pairs spanning vision, tabular, and language modalities, with varying label proportions. Unlike existing benchmarks, which primarily focus on shifts in class-conditional $p(x|y)$, our benchmark also focuses on label marginal shifts. First, we assess 13 popular domain adaptation methods, demonstrating more widespread failures under label proportion shifts than were previously known. Next, we develop an effective two-step meta-algorithm that is compatible with most domain adaptation heuristics: (i) pseudo-balance the data at each epoch; and (ii) adjust the final classifier with target label distribution estimate. The meta-algorithm improves existing domain adaptation heuristics under large label proportion shifts, often by 2--10\% accuracy points, while conferring minimal effect ($<$0.5\%) when label proportions do not shift. We hope that these findings and the availability of RLSbench will encourage researchers to rigorously evaluate proposed methods in relaxed label shift settings. Code is publicly available at https://github.com/acmi-lab/RLSbench.