Matching while Learning

Matching while Learning
复制标题

DOI:
10.1145/3033274.3084095
复制
发表时间:
2016-03
期刊:
Proceedings of the 2017 ACM Conference on Economics and Computation
影响因子:
--
通讯作者:
Ramesh Johari;Vijay Kamble;Yashodhan Kanoria
Ramesh Johari;Vijay Kamble;Yashodhan Kanoria
中科院分区:
其他
文献类型:
--
作者:
Ramesh Johari;Vijay Kamble;Yashodhan Kanoria

文献摘要

被引文献

相似文献

我们考虑了服务平台所面临的问题,该平台需要将供应与需求相匹配,但也需要了解新到达者的属性,以便在未来更好地匹配他们。我们引入了一个基准模型与异质工人和工作,随着时间的推移到达。作业类型对于平台是已知的,但工作者类型是未知的,必须通过观察匹配结果来学习。工人在完成一定数量的工作后离开。匹配的收益取决于两种类型,目标是最大化收益的稳态累积率。我们的主要贡献是一个完整的表征结构的最优策略的限制,每个工人执行许多工作。对于每个工人来说,平台都面临着一个权衡:短视地最大化收益(剥削)和学习工人的类型(探索)。这就产生了大量的多臂强盗问题,每个工人一个,通过对不同类型工作的可用性的约束(能力约束)耦合在一起。我们发现,平台应该估计每种工作类型的影子价格,并使用这些价格调整后的收益,首先,确定其学习目标,然后,对于每个工人,(i)在探索阶段平衡学习与收益,(ii)在开发阶段实现学习目标后进行近视匹配。
We consider the problem faced by a service platform that needs to match supply with demand but also to learn attributes of new arrivals in order to match them better in the future. We introduce a benchmark model with heterogeneous workers and jobs that arrive over time. Job types are known to the platform, but worker types are unknown and must be learned by observing match outcomes. Workers depart after performing a certain number of jobs. The payoff from a match depends on the pair of types and the goal is to maximize the steady-state rate of accumulation of payoff. Our main contribution is a complete characterization of the structure of the optimal policy in the limit that each worker performs many jobs. The platform faces a trade-off for each worker between myopically maximizing payoffs (exploitation) and learning the type of the worker (exploration). This creates a multitude of multi-armed bandit problems, one for each worker, coupled together by the constraint on the availability of jobs of different types (capacity constraints). We find that the platform should estimate a shadow price for each job type, and use the payoffs adjusted by these prices, first, to determine its learning goals and then, for each worker, (i) to balance learning with payoffs during the exploration phase, and (ii) to myopically match after it has achieved its learning goals during the exploitation phase.