Interpretable and Effective Reinforcement Learning for Attacking against Graph-based Rumor Detection

Interpretable and Effective Reinforcement Learning for Attacking against Graph-based Rumor Detection
复制标题

DOI:
10.1109/ijcnn54540.2023.10191290
复制
发表时间:
2022-01
期刊:
2023 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Yuefei Lyu;Xiaoyu Yang;Jiaxin Liu;Sihong Xie;Xi Zhang
Yuefei Lyu;Xiaoyu Yang;Jiaxin Liu;Sihong Xie;Xi Zhang
中科院分区:
其他
文献类型:
--
作者:
Yuefei Lyu;Xiaoyu Yang;Jiaxin Liu;Sihong Xie;Xi Zhang

文献摘要

相似文献

社交网络经常被谣言污染,这可以通过图神经网络等高级模型来检测。然而,这些模型容易受到攻击,发现和了解这些漏洞对于强大的谣言检测至关重要。为了发现细微的漏洞,我们设计了一种基于强化学习的攻击算法来伪装谣言对抗黑盒检测器。我们解决了指数级的大状态空间、高阶图依赖关系和排名依赖关系,这些对于问题设置来说是独一无二的,但对最先进的端到端方法来说是根本的挑战。我们设计了特定于领域的特征,这些特征对奖励具有因果影响,因此即使是线性策略也可以达到具有额外可解释性的强大攻击。为了加速策略优化,我们设计了一种信用分配方法,该方法将延迟和聚集的奖励按比例分解为原子攻击动作,以增强特征-奖励关联;(Ii)基于奖励方差分析和预测分布的贝叶斯分析,使用依赖于时间的控制变量来降低由于大状态动作空间和长攻击时间而导致的预测方差。在谣言检测任务的两个真实数据集上,我们证明了:(I)与基于规则的攻击方法和端到端攻击方法相比,学习的攻击策略在广泛的目标模型上的有效性;(Ii)所提出的信用分配策略和方差减少组件的有用性;(Iii)攻击策略的可解释性。
Social networks are frequently polluted by rumors, which can be detected by advanced models such as graph neural networks. However, the models are vulnerable to attacks, and discovering and understanding the vulnerabilities is critical to robust rumor detection. To discover subtle vulnerabilities, we design a attacking algorithm based on reinforcement learning to camouflage rumors against black-box detectors. We address exponentially large state spaces, high-order graph dependencies, and ranking dependencies, which are unique to the problem setting but fundamentally challenging for the state-of-the-art end-to-end approaches. We design domain-specific features that have causal effect on the reward, so that even a linear policy can arrive at powerful attacks with additional interpretability. To speed up policy optimization, we devise: (i) a credit assignment method that proportionally decomposes delayed and aggregated rewards to atomic attacking actions for enhance feature-reward associations; (ii) a time-dependent control variate to reduce prediction variance due to large state-action spaces and long attack horizon, based on reward variance analysis and a Bayesian analysis of the prediction distribution. On two real world datasets of rumor detection tasks, we demonstrate: (i) the effectiveness of the learned attacking policy on a wide spectrum of target models compared to both rule-based and end-to-end attacking approaches; (ii) the usefulness of the proposed credit assignment strategy and variance reduction components; (iii) the interpretability of the attacking policy.