On Tilted Losses in Machine Learning: Theory and Applications

On Tilted Losses in Machine Learning: Theory and Applications
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Tian Li;Ahmad Beirami;Maziar Sanjabi;Virginia Smith
Tian Li;Ahmad Beirami;Maziar Sanjabi;Virginia Smith
中科院分区:
其他
文献类型:
--
作者:
Tian Li;Ahmad Beirami;Maziar Sanjabi;Virginia Smith

文献摘要

被引文献

相似文献

指数倾斜是一种常用于统计学、概率论、信息论和优化等领域的技术,用于实现参数化的分布变换。尽管它在相关领域应用广泛,但在机器学习中尚未得到普遍应用。在这项研究中,我们旨在通过探索倾斜在风险最小化中的应用来弥合这一差距。我们研究了经验风险最小化(ERM)的一种简单扩展——倾斜经验风险最小化(TERM),它利用指数倾斜来灵活调整单个损失的影响。由此产生的框架具有几个有用的特性:我们证明TERM可以分别增加或减少异常值的影响,以实现公平性或稳健性;具有有助于泛化的方差缩减特性;并且可以看作是损失尾部概率的平滑近似。我们的工作在TERM与相关目标(如风险价值、条件风险价值和分布鲁棒优化(DRO))之间建立了严谨的联系。我们开发了用于求解TERM的批量和随机一阶优化方法,为求解器提供了收敛保证,并表明相对于常见的替代方法,该框架能够高效求解。最后,我们证明TERM可用于机器学习中的多种应用,例如在子群体间实施公平性、减轻异常值的影响以及处理类别不平衡问题。尽管TERM对传统ERM目标的修改很直接,但我们发现该框架始终优于ERM,并且与针对特定问题的最先进方法相比,也能展现出具有竞争力的性能。
Exponential tilting is a technique commonly used in fields such as statistics, probability, information theory, and optimization to create parametric distribution shifts. Despite its prevalence in related fields, tilting has not seen widespread use in machine learning. In this work, we aim to bridge this gap by exploring the use of tilting in risk minimization. We study a simple extension to ERM -- tilted empirical risk minimization (TERM) -- which uses exponential tilting to flexibly tune the impact of individual losses. The resulting framework has several useful properties: We show that TERM can increase or decrease the influence of outliers, respectively, to enable fairness or robustness; has variance-reduction properties that can benefit generalization; and can be viewed as a smooth approximation to the tail probability of losses. Our work makes rigorous connections between TERM and related objectives, such as Value-at-Risk, Conditional Value-at-Risk, and distributionally robust optimization (DRO). We develop batch and stochastic first-order optimization methods for solving TERM, provide convergence guarantees for the solvers, and show that the framework can be efficiently solved relative to common alternatives. Finally, we demonstrate that TERM can be used for a multitude of applications in machine learning, such as enforcing fairness between subgroups, mitigating the effect of outliers, and handling class imbalance. Despite the straightforward modification TERM makes to traditional ERM objectives, we find that the framework can consistently outperform ERM and deliver competitive performance with state-of-the-art, problem-specific approaches.