Towards Return Parity in Markov Decision Processes

Towards Return Parity in Markov Decision Processes
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
--
影响因子:
--
通讯作者:
Jianfeng Chi;Jian Shen;Xinyi Dai;Weinan Zhang;Yuan Tian;Han Zhao
Jianfeng Chi;Jian Shen;Xinyi Dai;Weinan Zhang;Yuan Tian;Han Zhao
中科院分区:
其他
文献类型:
--
作者:
Jianfeng Chi;Jian Shen;Xinyi Dai;Weinan Zhang;Yuan Tian;Han Zhao

文献摘要

被引文献

相似文献

机器学习模型在高风险领域做出的决策可能会随着时间的推移产生持久的影响。然而,标准公平准则在时域上的静态设置中的幼稚应用可能导致延迟和不利影响。为了理解性能差异的动态,我们研究了马尔可夫决策过程(MDP)中的公平性问题。具体来说,我们提出了回报平价,一个公平的概念,需要来自不同人口群体的MDP共享相同的状态和动作空间,以实现大致相同的预期时间折扣奖励。我们首先提供了一个分解定理的回报差距,它分解的回报差距的任何两个共享相同的状态和动作空间的MDP到组明智的奖励函数之间的距离,组策略的差异,和状态访问分布之间的差异引起的组策略。出于我们的分解定理,我们提出了算法,以减轻回报差距,通过学习一个共享的组策略与状态访问分布对齐使用积分概率度量。我们进行实验来证实我们的结果,表明该算法可以成功地关闭差距,同时保持两个现实世界的推荐系统基准数据集的政策的性能。
Algorithmic decisions made by machine learning models in high-stakes domains may have lasting impacts over time. However, naive applications of standard fairness criterion in static settings over temporal domains may lead to delayed and adverse effects. To understand the dynamics of performance disparity, we study a fairness problem in Markov decision processes (MDPs). Specifically, we propose return parity, a fairness notion that requires MDPs from different demographic groups that share the same state and action spaces to achieve approximately the same expected time-discounted rewards. We first provide a decomposition theorem for return disparity, which decomposes the return disparity of any two MDPs sharing the same state and action spaces into the distance between group-wise reward functions, the discrepancy of group policies, and the discrepancy between state visitation distributions induced by the group policies. Motivated by our decomposition theorem, we propose algorithms to mitigate return disparity via learning a shared group policy with state visitation distributional alignment using integral probability metrics. We conduct experiments to corroborate our results, showing that the proposed algorithm can successfully close the disparity gap while maintaining the performance of policies on two real-world recommender system benchmark datasets.