Using Contextual Bandits with Behavioral Constraints for Constrained Online Movie Recommendation

Using Contextual Bandits with Behavioral Constraints for Constrained Online Movie Recommendation
复制标题

DOI:
10.24963/ijcai.2018/843
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
Avinash Balakrishnan;Djallel Bouneffouf;Nicholas Mattei;F. Rossi
Avinash Balakrishnan;Djallel Bouneffouf;Nicholas Mattei;F. Rossi
中科院分区:
其他
文献类型:
--
作者:
Avinash Balakrishnan;Djallel Bouneffouf;Nicholas Mattei;F. Rossi

文献摘要

被引文献

相似文献

人工智能系统通过对其所采取行动的奖励反馈来学习,越来越多地部署在对我们的日常生活具有重大影响的领域。在许多情况下,奖励不应该是唯一的指导标准,因为法规、价值观、偏好或道德原则还强加了额外的限制和/或优先事项。我们详细介绍了一个新颖的在线系统,该系统基于上下文盗贼框架的扩展,通过观察学习一组行为约束,并在在线环境中做出决策时使用这些约束作为指导,同时仍然对奖励反馈做出反应。此外,我们的系统可以突出上下文的特征,这些特征被更多地预测为更有价值和/或符合行为约束。我们通过构建一个在线电影推荐代理的交互界面来演示该系统,并表明我们的系统能够在不显著降低整体性能的情况下在一组行为约束内行动。
AI systems that learn through reward feedback about the actions they take are increasingly deployed in domains that have significant impact on our daily life. In many cases the rewards should not be the only guiding criteria, as there are additional constraints and/or priorities imposed by regulations, values, preferences, or ethical principles. We detail a novel online system, based on an extension of the contextual bandits framework, that learns a set of behavioral constraints by observation and uses these constraints as a guide when making decisions in an online setting while still being reactive to reward feedback. In addition, our system can highlight features of the context which are more predicted to be more rewarding and/or are in line with the behavioral constraints. We demonstrate the system by building an interactive interface for an online movie recommendation agent and show that our system is able to act within a set of behavior constraints without significantly degrading overall performance.