Label-Imbalanced and Group-Sensitive Classification under Overparameterization

Label-Imbalanced and Group-Sensitive Classification under Overparameterization
复制标题

DOI:
--
复制
发表时间:
2021-03
期刊:
--
影响因子:
--
通讯作者:
Ganesh Ramachandra Kini;Orestis Paraskevas;Samet Oymak;Christos Thrampoulidis
Ganesh Ramachandra Kini;Orestis Paraskevas;Samet Oymak;Christos Thrampoulidis
中科院分区:
其他
文献类型:
--
作者:
Ganesh Ramachandra Kini;Orestis Paraskevas;Samet Oymak;Christos Thrampoulidis

文献摘要

相似文献

标签不平衡和组敏感分类的目标是优化相关指标,如平衡错误和平等机会。经典的方法,如加权交叉熵,在训练深度网络到训练的终端阶段(TPT)时失败,即训练超过零训练误差。这一观察结果激发了最近的一系列活动,即根据促进少数群体获得更大幅度的直观机制,开发启发式替代方案。与之前的分析相反,我们遵循原则性分析,解释不同的损失调整如何影响利润率。首先,我们证明了所有的线性分类训练TPT,有必要引入乘法,而不是添加剂,logit调整,使类间的利润率适当地改变。为了证明这一点,我们发现了一个连接的乘法CE修改的成本敏感的支持向量机。也许与直觉相反,我们还发现,在训练开始时,相同的乘法权重实际上会损害少数类。因此,虽然添加剂的调整是无效的TPT,我们表明,他们可以通过对抗乘法权重的初始负面影响,加快收敛。出于这些发现,我们制定的矢量缩放(VS)的损失,捕获现有的技术作为特殊情况。此外,我们引入了一个自然扩展的VS损失组敏感的分类,从而治疗两种常见类型的不平衡(标签/组)在一个统一的方式。重要的是,我们在最先进的数据集上的实验与我们的理论见解完全一致,并证实了我们算法的上级性能。最后,对于不平衡的高斯混合数据,我们进行了泛化分析,揭示了平衡/标准误差和平等机会之间的权衡。
The goal in label-imbalanced and group-sensitive classification is to optimize relevant metrics such as balanced error and equal opportunity. Classical methods, such as weighted cross-entropy, fail when training deep nets to the terminal phase of training (TPT), that is training beyond zero training error. This observation has motivated recent flurry of activity in developing heuristic alternatives following the intuitive mechanism of promoting larger margin for minorities. In contrast to previous heuristics, we follow a principled analysis explaining how different loss adjustments affect margins. First, we prove that for all linear classifiers trained in TPT, it is necessary to introduce multiplicative, rather than additive, logit adjustments so that the interclass margins change appropriately. To show this, we discover a connection of the multiplicative CE modification to the cost-sensitive support-vector machines. Perhaps counterintuitively, we also find that, at the start of training, the same multiplicative weights can actually harm the minority classes. Thus, while additive adjustments are ineffective in the TPT, we show that they can speed up convergence by countering the initial negative effect of the multiplicative weights. Motivated by these findings, we formulate the vector-scaling (VS) loss, that captures existing techniques as special cases. Moreover, we introduce a natural extension of the VS-loss to group-sensitive classification, thus treating the two common types of imbalances (label/group) in a unifying way. Importantly, our experiments on state-of-the-art datasets are fully consistent with our theoretical insights and confirm the superior performance of our algorithms. Finally, for imbalanced Gaussian-mixtures data, we perform a generalization analysis, revealing tradeoffs between balanced / standard error and equal opportunity.