Learn to Match with No Regret: Reinforcement Learning in Markov Matching Markets

Learn to Match with No Regret: Reinforcement Learning in Markov Matching Markets
复制标题

DOI:
10.48550/arxiv.2203.03684
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Yifei Min;Tianhao Wang;Ruitu Xu;Zhaoran Wang;Michael I. Jordan;Zhuoran Yang
Yifei Min;Tianhao Wang;Ruitu Xu;Zhaoran Wang;Michael I. Jordan;Zhuoran Yang
中科院分区:
其他
文献类型:
--
作者:
Yifei Min;Tianhao Wang;Ruitu Xu;Zhaoran Wang;Michael I. Jordan;Zhuoran Yang

文献摘要

被引文献

相似文献

我们研究了一个马尔可夫匹配市场,涉及市场两侧的规划者和一组策略代理。在每个步骤中,代理都会看到一个动态上下文,其中上下文决定了效用。规划者控制环境的转变以最大化累积社会福利,而智能体的目标是在每一步找到近视的稳定匹配。这样的设置涵盖了包括乘车共享平台在内的一系列应用程序。我们通过提出一个将乐观值迭代与最大权重匹配相结合的强化学习框架来形式化该问题。所提出的算法解决了顺序探索、匹配稳定性和函数逼近的耦合挑战。我们证明该算法实现了亚线性后悔。
We study a Markov matching market involving a planner and a set of strategic agents on the two sides of the market. At each step, the agents are presented with a dynamical context, where the contexts determine the utilities. The planner controls the transition of the contexts to maximize the cumulative social welfare, while the agents aim to find a myopic stable matching at each step. Such a setting captures a range of applications including ridesharing platforms. We formalize the problem by proposing a reinforcement learning framework that integrates optimistic value iteration with maximum weight matching. The proposed algorithm addresses the coupled challenges of sequential exploration, matching stability, and function approximation. We prove that the algorithm achieves sublinear regret.