User Tampering in Reinforcement Learning Recommender Systems

User Tampering in Reinforcement Learning Recommender Systems
复制标题

强化学习推荐系统中的用户篡改

DOI:
--
复制
发表时间:
2021
期刊:
AAAI/ACM Conference on AI, Ethics, and Society
影响因子:
--
通讯作者:
Charles Evans
Charles Evans
中科院分区:
--
文献类型:
--
作者:
Atoosa Kasirzadeh;Charles Evans

文献摘要

被引文献

相似文献

在本文中,我们引入了新的形式化方法,并提供了经验证据来强调基于强化学习(RL)的推荐算法中普遍存在的一个独特的安全问题-“用户篡改”。用户篡改是一种基于RL的推荐系统可以通过媒体用户的建议来操纵媒体用户的意见作为策略的一部分以最大化长期用户参与的情况。我们使用正式的因果建模技术,批判性地分析文献中提出的解决方案,实现可扩展的基于RL的推荐系统,我们观察到,这些方法不足以防止用户篡改。此外,我们评估现有的缓解策略奖励篡改的问题,并表明这些方法是不够的,在解决用户篡改的背景下,建议的独特现象。我们进一步加强了我们的研究结果与模拟研究的RL为基础的推荐系统,专注于传播的政治内容。我们的研究表明,Q学习算法始终学习利用其机会,利用其早期建议来验证模拟用户,以便在与该诱导极化一致的后续建议中取得更一致的成功。我们的研究结果强调了开发更安全的基于RL的推荐系统的必要性,并建议实现这种安全性将需要从设计上的根本转变,远离我们在最近的文献中看到的方法。
In this paper, we introduce new formal methods and provide empirical evidence to highlight a unique safety concern prevalent in reinforcement learning (RL)-based recommendation algorithms – ’user tampering.’ User tampering is a situation where an RL-based recommender system may manipulate a media user’s opinions through its suggestions as part of a policy to maximize long-term user engagement. We use formal techniques from causal modeling to critically analyze prevailing solutions proposed in the literature for implementing scalable RL-based recommendation systems, and we observe that these methods do not adequately prevent user tampering. Moreover, we evaluate existing mitigation strategies for reward tampering issues, and show that these methods are insufficient in addressing the distinct phenomenon of user tampering within the context of recommendations. We further reinforce our findings with a simulation study of an RL-based recommendation system focused on the dissemination of political content. Our study shows that a Q-learning algorithm consistently learns to exploit its opportunities to polarize simulated users with its early recommendations in order to have more consistent success with subsequent recommendations that align with this induced polarization. Our findings emphasize the necessity for developing safer RL-based recommendation systems and suggest that achieving such safety would require a fundamental shift in the design away from the approaches we have seen in the recent literature.