A PCA-Based Model to Predict Adversarial Examples on Q-Learning of Path Finding

A PCA-Based Model to Predict Adversarial Examples on Q-Learning of Path Finding
复制标题

基于 PCA 的路径查找 Q 学习对抗样本预测模型

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Data Science in Cyberspace
影响因子:
--
通讯作者:
Zhen Han
Zhen Han
中科院分区:
--
文献类型:
--
作者:
Yingxiao Xiang;Wenjia Niu;Jiqiang Liu;Tong Chen;Zhen Han

文献摘要

被引文献

相似文献

强化学习是人工智能中机器学习的一种核心学习方法,已被广泛应用于机器人自动寻路模块的实现。然而,最近的研究表明,设计的对抗性示例可以对Atari游戏中的强化学习进行算法级攻击。显然,这种攻击也可能被引入到基于强化学习的自动路径查找中。在本文中,我们专注于对抗性的基于示例的攻击的一个代表性的强化学习命名为Q学习在自动路径查找。我们提出了一种基于影响因素和相应权重的概率输出模型来预测对抗性样本。通过对能量点引力、关键点引力、路径引力、夹角和平静点5个因素的计算,构建了一个自然线性模型来拟合这些因素,并基于主成分分析(PCA)计算权重参数。通过大量的实验,我们成功地发现了对抗性的例子,第一次在Q-学习的路径发现和我们的模型可以做出令人满意的预测。在保证查全率的前提下,通过适当的参数设置,该模型的查准率可以达到70%。
Reinforcement learning is a core learning approach of machine learning in artificial intelligence, which has been widely used to realize the module of automatic path finding for robots. However, recent researches have shown that the designed adversarial examples can make an algorithm-level attack on reinforcement learning in the Atari game. Obviously, such attack may be also brought into the automatic path finding based on reinforcement learning. In this paper, we focus on the adversarial example-based attack on a representative reinforcement learning named Q-learning in automatic path finding. We propose a probabilistic output model based on the influence factors and the corresponding weights to predict the adversarial examples. Through calculation on five factors including the energy point gravitation, the key point gravitation, the path gravitation, the included angle and the placid point, a natural linear model is constructed to fit these factors with the weight parameters computation based on the principal component analysis(PCA). Through massive experiments, we successfully find the adversarial examples for the first time on Q-learning in path finding and our model can make a satisfactory prediction. Under a guaranteed recall, the precision of the proposed model can reach to 70% with the proper parameter setting.