Policy Optimization with Advantage Regularization for Long-Term Fairness in Decision Systems

Policy Optimization with Advantage Regularization for Long-Term Fairness in Decision Systems
复制标题

DOI:
10.48550/arxiv.2210.12546
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Eric Yang Yu;Zhizhen Qin;Min Kyung Lee;Sicun Gao
Eric Yang Yu;Zhizhen Qin;Min Kyung Lee;Sicun Gao
中科院分区:
其他
文献类型:
--
作者:
Eric Yang Yu;Zhizhen Qin;Min Kyung Lee;Sicun Gao

文献摘要

被引文献

相似文献

长期公平性是在高风险决策环境中设计和部署基于学习的决策系统时考虑的一个重要因素。最近的工作提出了使用马尔可夫决策过程(MDP)制定决策与长期的公平性要求,在动态变化的环境中,并表现出直接部署启发式和基于规则的政策,在静态环境中工作良好的主要挑战。我们表明,深度强化学习的策略优化方法可用于找到严格更好的决策策略,与先前已知的策略相比,这些策略通常可以实现更高的整体效用和更少的公平性要求。特别是,我们提出了新的方法,在政策优化的公平性要求,通过正规化的优势评估不同的行动。我们提出的方法可以很容易地施加公平性约束,而无需奖励工程或牺牲训练效率。我们在三个已建立的案例研究中进行了详细的分析,包括事件监测中的注意力分配,银行贷款审批和人口网络中的疫苗分配。
Long-term fairness is an important factor of consideration in designing and deploying learning-based decision systems in high-stake decision-making contexts. Recent work has proposed the use of Markov Decision Processes (MDPs) to formulate decision-making with long-term fairness requirements in dynamically changing environments, and demonstrated major challenges in directly deploying heuristic and rule-based policies that worked well in static environments. We show that policy optimization methods from deep reinforcement learning can be used to find strictly better decision policies that can often achieve both higher overall utility and less violation of the fairness requirements, compared to previously-known strategies. In particular, we propose new methods for imposing fairness requirements in policy optimization by regularizing the advantage evaluation of different actions. Our proposed methods make it easy to impose fairness constraints without reward engineering or sacrificing training efficiency. We perform detailed analyses in three established case studies, including attention allocation in incident monitoring, bank loan approval, and vaccine distribution in population networks.