The Two Dimensions of Worst-case Training and Their Integrated Effect for Out-of-domain Generalization

The Two Dimensions of Worst-case Training and Their Integrated Effect for Out-of-domain Generalization
复制标题

DOI:
10.1109/cvpr52688.2022.00941
复制
发表时间:
2022-04
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Zeyi Huang;Haohan Wang;Dong Huang;Yong Jae Lee;Eric P. Xing
Zeyi Huang;Haohan Wang;Dong Huang;Yong Jae Lee;Eric P. Xing
中科院分区:
其他
文献类型:
--
作者:
Zeyi Huang;Haohan Wang;Dong Huang;Yong Jae Lee;Eric P. Xing

文献摘要

相似文献

事实证明,强调数据“难学”部分的训练是提高机器学习模型泛化能力的有效方法,特别是在重视鲁棒性(例如跨分布泛化)的环境中。现有文献讨论这个“难学”概念主要是沿着样本维度或特征维度展开。在本文中,我们的目标是引入一个合并这两个维度的简单视图,通过强调样本和特征维度的最坏情况,产生一种新的、简单而有效的启发式训练机器学习模型。我们根据“沿二维的最坏情况”的概念将我们的方法命名为 W2D。我们验证了这个想法,并证明了其相对于标准基准的实证强度。
Training with an emphasis on “hard-to-learn” components of the data has been proven as an effective method to improve the generalization of machine learning models, especially in the settings where robustness (e.g., generalization across distributions) is valued. Existing literature discussing this “hard-to-learn” concept are mainly expanded either along the dimension of the samples or the dimension of the features. In this paper, we aim to introduce a simple view merging these two dimensions, leading to a new, simple yet effective, heuristic to train machine learning models by emphasizing the worst-cases on both the sample and the feature dimensions. We name our method W2D following the concept of “Worst-case along Two Dimensions”. We validate the idea and demonstrate its empirical strength over standard benchmarks.