Sample Selection for Fair and Robust Training

Sample Selection for Fair and Robust Training
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
Yuji Roh;Kangwook Lee;Steven Euijong Whang;Changho Suh
Yuji Roh;Kangwook Lee;Steven Euijong Whang;Changho Suh
中科院分区:
其他
文献类型:
--
作者:
Yuji Roh;Kangwook Lee;Steven Euijong Whang;Changho Suh

文献摘要

相似文献

公平性和鲁棒性是值得信赖的AI的关键要素,需要一起解决。公平性是关于学习无偏模型,而鲁棒性是关于从损坏的数据中学习,并且已知仅解决其中一个可能会对另一个产生不利影响。在这项工作中,我们提出了一个基于样本选择的算法,公平和强大的训练。为此,我们制定了一个组合优化问题的无偏选择的样本中存在的数据损坏。观察到解决这个优化问题是强NP难的,我们提出了一个贪婪算法,是有效的,在实践中。实验表明,我们的算法获得的公平性和鲁棒性,优于或相当于国家的最先进的技术,无论是在合成和基准真实的数据集。此外,与其他公平和强大的训练基线不同,我们的算法可以通过仅修改批次选择中的采样步骤来使用,而无需更改训练算法或利用额外的干净数据。
Fairness and robustness are critical elements of Trustworthy AI that need to be addressed together. Fairness is about learning an unbiased model while robustness is about learning from corrupted data, and it is known that addressing only one of them may have an adverse affect on the other. In this work, we propose a sample selection-based algorithm for fair and robust training. To this end, we formulate a combinatorial optimization problem for the unbiased selection of samples in the presence of data corruption. Observing that solving this optimization problem is strongly NP-hard, we propose a greedy algorithm that is efficient and effective in practice. Experiments show that our algorithm obtains fairness and robustness that are better than or comparable to the state-of-the-art technique, both on synthetic and benchmark real datasets. Moreover, unlike other fair and robust training baselines, our algorithm can be used by only modifying the sampling step in batch selection without changing the training algorithm or leveraging additional clean data.