A Deep Reinforcement Learning Recommender System With Multiple Policies for Recommendations

A Deep Reinforcement Learning Recommender System With Multiple Policies for Recommendations
复制标题

DOI:
10.1109/tii.2022.3209290
复制
发表时间:
2023-02
影响因子:
12.3
通讯作者:
Mingsheng Fu;Liwei Huang;Ananya Rao;Athirai Aravazhi Irissappane;Jie Zhang;Hong Qu
Mingsheng Fu;Liwei Huang;Ananya Rao;Athirai Aravazhi Irissappane;Jie Zhang;Hong Qu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Mingsheng Fu;Liwei Huang;Ananya Rao;Athirai Aravazhi Irissappane;Jie Zhang;Hong Qu

文献摘要

相似文献

基于深度强化学习(DRL)的推荐系统能够逐步捕获用户偏好,适用于用户冷启动问题。然而,大多数现有的基于DRL的推荐系统都是次优的,因为它们使用相同的策略来适应不同用户的动态。我们将推荐重新定义为一个多任务马尔可夫决策过程,其中每个任务代表一组相似的用户。由于相似的用户具有更紧密的动态,因此特定于任务的策略比适用于所有用户的单一通用策略更有效。为了向冷启动用户提供建议,我们使用默认策略来收集一些初始交互来识别用户任务,然后使用特定于任务的策略。我们使用Q-学习来优化框架,并通过与任务相关的互信息来考虑任务的不确定性。在三个真实世界的数据集上进行了实验,以验证我们所提出的框架的有效性。
Deep reinforcement learning (DRL) based recommender systems are suitable for user cold-start problems as they can capture user preferences progressively. However, most existing DRL-based recommender systems are suboptimal, since they use the same policy to suit the dynamics of different users. We reformulate recommendation as a multitask Markov Decision Process, where each task represents a set of similar users. Since similar users have closer dynamics, a task-specific policy is more effective than a single universal policy for all users. To make recommendations for cold-start users, we use a default policy to collect some initial interactions to identify the user task, after which a task-specific policy is employed. We use Q-learning to optimize our framework and consider the task uncertainty by the mutual information regarding tasks. Experiments are conducted on three real-world datasets to verify the effectiveness of our proposed framework.