Distributionally Robust Losses for Latent Covariate Mixtures

Distributionally Robust Losses for Latent Covariate Mixtures
复制标题

DOI:
10.1287/opre.2022.2363
复制
发表时间:
2020-07
期刊:
Oper. Res.
影响因子:
--
通讯作者:
John C. Duchi;Tatsunori B. Hashimoto;Hongseok Namkoong
John C. Duchi;Tatsunori B. Hashimoto;Hongseok Namkoong
中科院分区:
其他
文献类型:
--
作者:
John C. Duchi;Tatsunori B. Hashimoto;Hongseok Namkoong

文献摘要

被引文献

相似文献

可靠的机器学习用于训练机器学习的结构性优化数据集(ML)模型通常会遭受采样偏见和占主导地位的分数,以优化平均性能,并在“分配非常强大”中表现不佳潜在协变量混合物的损失,“约翰·杜奇(John Duchi),Tatsunori Hashimoto和Hongseok Namkoong为训练ML模型制定了均匀的训练ML模型。案例亚群的性能并提供有限样本(非参数)的融合。相似性,葡萄酒质量和累犯预测任务,并观察到在看不见的亚群体中的性能显着提高。
Reliable Machine Learning via Structured Distributionally Robust Optimization Data sets used to train machine learning (ML) models often suffer from sampling biases and underrepresent marginalized groups. Standard machine learning models are trained to optimize average performance and perform poorly on tail subpopulations. In “Distributionally Robust Losses for Latent Covariate Mixtures,” John Duchi, Tatsunori Hashimoto, and Hongseok Namkoong formulate a DRO approach for training ML models to perform uniformly well over subpopulations. They design a worst case optimization procedure over structured distribution shifts salient in predictive applications: shifts in (a subset of) covariates. The authors propose a convex procedure that controls worst case subpopulation performance and provide finite-sample (nonparametric) convergence guarantees. Empirically, they demonstrate their worst case procedure on lexical similarity, wine quality, and recidivism prediction tasks and observe significantly improved performance across unseen subpopulations.