Deep Reinforcement Learning Based Dynamic Reputation Policy in 5G Based Vehicular Communication Networks

Deep Reinforcement Learning Based Dynamic Reputation Policy in 5G Based Vehicular Communication Networks
复制标题

DOI:
10.1109/tvt.2021.3079379
复制
发表时间:
2021-06
影响因子:
6.8
通讯作者:
Sohan Gyawali;Y. Qian;R. Hu
Sohan Gyawali;Y. Qian;R. Hu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Sohan Gyawali;Y. Qian;R. Hu

文献摘要

相似文献

车辆网络很容易受到网络内恶意车辆或基础设施的各种攻击。协作不当行为检测系统可用于检测这些内部或内部攻击。然而,在协作式不当行为检测系统中,攻击者可能会通过发送错误反馈来降低检测精度。信任模型可用于刺激车辆发送真实的反馈。但是,攻击者可以利用弱或强的信誉更新方法。动态信任或声誉更新策略可用于刺激车辆发送真实反馈。在本文中,我们提出了一种基于深度强化学习的动态声誉更新策略。在所提出的方案中,使用 Dempster-Shafer 理论在车辆边缘计算 (VEC) 服务器中组合来自车辆的反馈,并将结果用于预测真实消息的平均数量。然后,VEC 使用深度强化学习来确定最佳声誉更新策略,以刺激车辆发送真实反馈。此外,通过广泛的模拟,我们表明,与现有的声誉更新策略相比,所提出的动态声誉策略在真实反馈的平均数量方面更好。
Vehicular networks are vulnerable to various attacks from malicious vehicles or infrastructures within a network. The collaborative misbehavior detection system can be used to detect these internal or insider attacks. However, in a collaborative misbehavior detection system, an attacker may lower the detection accuracy by sending false feedback. A trust model can be used to stimulate vehicles to send true feedbacks. However, an attacker can take advantage of weak or strong reputation update methods. A dynamic trust or reputation update policy can be used to stimulate vehicles to send true feedbacks. In this paper, we propose a deep reinforcement learning based dynamic reputation update policy. In the proposed scheme, feedbacks from vehicles are combined in vehicular edge computing (VEC) servers using Dempster-Shafer theory and the results are used to predict the average number of true messages. VEC then uses deep reinforcement learning to determine the optimum reputation update policy to stimulate vehicles to send true feedbacks. In addition, through extensive simulations, we show that the proposed dynamic reputation-policy is better in terms of the average number of true feedbacks compared to the existing reputation update policy.