Learning Personalized Product Recommendations with Customer Disengagement

Learning Personalized Product Recommendations with Customer Disengagement
复制标题

通过客户脱离学习个性化产品推荐

DOI:
--
复制
发表时间:
2018
期刊:
Manufacturing & Service Operations Management
影响因子:
--
通讯作者:
Divya Singhvi
Divya Singhvi
中科院分区:
--
文献类型:
--
作者:
Hamsa Bastani;P. Harsha;G. Perakis;Divya Singhvi

文献摘要

参考文献

被引文献

相似文献

问题定义:我们研究在客户偏好未知的情况下,平台上的个性化产品推荐。重要的是,当提供糟糕的推荐时,客户可能会退出。学术/实践相关性:在线平台通常使用强盗算法来个性化产品推荐,这平衡了探索与利用之间的权衡。然而,客户脱离——平台在实践中的一个显著特征——带来了一个新的挑战,因为探索可能导致客户放弃平台。我们提出了一种约束探索以提高性能的新算法。方法:我们使用一家大型航空公司广告活动的数据来展示客户脱离的证据;这激发了我们的脱离模式,在这种模式下,当提供不相关的建议时,客户可能会放弃平台。我们将客户偏好学习问题表述为广义线性强盗,显著的区别是客户的视界长度是过去推荐的函数。结果:我们证明了没有一种算法可以保持所有客户的参与。不幸的是,经典的强盗算法可能会过度探索,导致每个客户最终都退出。基于问题标量实例中最优策略的结构特性,我们提出通过使用整数程序预先约束动作空间来修改强盗学习策略。我们证明,这个简单的修改可以让我们的算法通过保持相当一部分客户的参与而表现良好。管理启示:如果客户有很高的脱离倾向,平台在了解客户偏好时应该小心避免过度探索。对电影推荐数据的数值实验表明,我们的算法可以显著提高客户粘性。
Problem definition: We study personalized product recommendations on platforms when customers have unknown preferences. Importantly, customers may disengage when offered poor recommendations. Academic/practical relevance: Online platforms often personalize product recommendations using bandit algorithms, which balance an exploration-exploitation trade-off. However, customer disengagement—a salient feature of platforms in practice—introduces a novel challenge because exploration may cause customers to abandon the platform. We propose a novel algorithm that constrains exploration to improve performance. Methodology: We present evidence of customer disengagement using data from a major airline’s ad campaign; this motivates our model of disengagement, where a customer may abandon the platform when offered irrelevant recommendations. We formulate the customer preference learning problem as a generalized linear bandit, with the notable difference that the customer’s horizon length is a function of past recommendations. Results: We prove that no algorithm can keep all customers engaged. Unfortunately, classical bandit algorithms provably overexplore, causing every customer to eventually disengage. Motivated by the structural properties of the optimal policy in a scalar instance of our problem, we propose modifying bandit learning strategies by constraining the action space up front using an integer program. We prove that this simple modification allows our algorithm to perform well by keeping a significant fraction of customers engaged. Managerial implications: Platforms should be careful to avoid overexploration when learning customer preferences if customers have a high propensity for disengagement. Numerical experiments on movie recommendations data demonstrate that our algorithm can significantly improve customer engagement.
DOI: 10.1109/msp.2018.2821706
发表时间: 2018-07-01
影响因子: 14.9
作者:
Chen, Yudong;Chi, Yuejie
通讯作者: Chi, Yuejie