RecoGym: A Reinforcement Learning Environment for the problem of Product Recommendation in Online Advertising

RecoGym: A Reinforcement Learning Environment for the problem of Product Recommendation in Online Advertising
复制标题

RecoGym:针对在线广告中产品推荐问题的强化学习环境

DOI:
--
复制
发表时间:
2018
期刊:
arXiv.org
影响因子:
--
通讯作者:
Alexandros Karatzoglou
Alexandros Karatzoglou
中科院分区:
--
文献类型:
--
作者:
D. Rohde;Stephen Bonner;Travis Dunlop;Flavian Vasile;Alexandros Karatzoglou

文献摘要

被引文献

相似文献

推荐系统在许多环境中变得无处不在,并采取多种形式,从电子商务商店中的产品推荐,到搜索引擎中的查询建议,再到社交网络中的好友推荐。当前主要基于历史数据监督学习的研究方向似乎显示出收益递减,许多从业者报告称监督学习的离线指标的改进与新提出的模型的在线性能之间存在差异。一个可能的原因是我们使用了错误的范式:在考虑收集历史性能数据的长期周期时,创建新版本的推荐模型,对其进行 A/B 测试,然后推出。我们看到强化学习(RL)设置有很多共同点,其中代理观察环境并对其采取行动,以便将其状态改变为更好的状态(具有更高奖励的状态)。为此,我们引入了 RecoGym,这是一种用于推荐的 RL 环境,它是由电子商务上的用户流量模式模型以及用户对发布商网站上的推荐的响应来定义的。我们相信,这是推荐系统研究领域向前迈出的重要一步,它可以开辟推荐系统和强化学习社区之间的合作途径,并导致离线和在线性能指标之间更好的协调。
Recommender Systems are becoming ubiquitous in many settings and take many forms, from product recommendation in e-commerce stores, to query suggestions in search engines, to friend recommendation in social networks. Current research directions which are largely based upon supervised learning from historical data appear to be showing diminishing returns with a lot of practitioners report a discrepancy between improvements in offline metrics for supervised learning and the online performance of the newly proposed models. One possible reason is that we are using the wrong paradigm: when looking at the long-term cycle of collecting historical performance data, creating a new version of the recommendation model, A/B testing it and then rolling it out. We see that there a lot of commonalities with the reinforcement learning (RL) setup, where the agent observes the environment and acts upon it in order to change its state towards better states (states with higher rewards). To this end we introduce RecoGym, an RL environment for recommendation, which is defined by a model of user traffic patterns on e-commerce and the users response to recommendations on the publisher websites. We believe that this is an important step forward for the field of recommendation systems research, that could open up an avenue of collaboration between the recommender systems and reinforcement learning communities and lead to better alignment between offline and online performance metrics.