Data Poisoning Attack against Recommender System Using Incomplete and Perturbed Data

Data Poisoning Attack against Recommender System Using Incomplete and Perturbed Data
复制标题

DOI:
10.1145/3447548.3467233
复制
发表时间:
2021-08
期刊:
Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Hengtong Zhang;Changxin Tian;Yaliang Li;Lu Su;Nan Yang;Wayne Xin Zhao;Jing Gao
Hengtong Zhang;Changxin Tian;Yaliang Li;Lu Su;Nan Yang;Wayne Xin Zhao;Jing Gao
中科院分区:
其他
文献类型:
--
作者:
Hengtong Zhang;Changxin Tian;Yaliang Li;Lu Su;Nan Yang;Wayne Xin Zhao;Jing Gao

文献摘要

相似文献

最近的研究表明,推荐系统由于其开放性而容易受到数据中毒攻击。在数据中毒攻击中,攻击者通常招募一组受控用户,将精心设计的用户-项目交互数据注入推荐模型的训练集中,以根据需要修改模型参数。因此,现有的攻击方法通常需要完全访问训练数据,以推断项目的特征并为受控用户制作虚假交互。然而,由于攻击者的数据收集能力有限以及对训练数据的访问受到限制,这种攻击方法在实践中可能并不可行,有时甚至受到服务提供商的隐私保护机制的干扰。这种设计与现实的差距可能会导致攻击失败。在本文中,我们通过提出两种新颖的对抗性攻击方法来处理用户-项目交互数据中的不完整性和扰动来填补这一空白。首先,我们提出了一个双层优化框架,该框架结合了概率生成模型来查找交互数据充足且未受到显着干扰的用户和项目,并利用这些用户和项目的数据来制作虚假的用户-项目交互。此外,我们逆转了推荐模型的学习过程,并开发了一种简单而有效的方法,可以结合特定于上下文的启发式规则来处理数据不完整性和扰动。针对三个代表性推荐模型对两个数据集进行的广泛实验表明,所提出的方法可以比现有方法实现更好的攻击性能。
Recent studies reveal that recommender systems are vulnerable to data poisoning attack due to their openness nature. In data poisoning attack, the attacker typically recruits a group of controlled users to inject well-crafted user-item interaction data into the recommendation model's training set to modify the model parameters as desired. Thus, existing attack approaches usually require full access to the training data to infer items' characteristics and craft the fake interactions for controlled users. However, such attack approaches may not be feasible in practice due to the attacker's limited data collection capability and the restricted access to the training data, which sometimes are even perturbed by the privacy preserving mechanism of the service providers. Such design-reality gap may cause failure of attacks. In this paper, we fill the gap by proposing two novel adversarial attack approaches to handle the incompleteness and perturbations in user-item interaction data. First, we propose a bi-level optimization framework that incorporates a probabilistic generative model to find the users and items whose interaction data is sufficient and has not been significantly perturbed, and leverage these users and items' data to craft fake user-item interactions. Moreover, we reverse the learning process of recommendation models and develop a simple yet effective approach that can incorporate context-specific heuristic rules to handle data incompleteness and perturbations. Extensive experiments on two datasets against three representative recommendation models show that the proposed approaches can achieve better attack performance than existing approaches.