Privacy-Preserving Federated Deep Reinforcement Learning for Mobility-as-a-Service

Privacy-Preserving Federated Deep Reinforcement Learning for Mobility-as-a-Service
复制标题

DOI:
10.1109/tits.2023.3317358
复制
发表时间:
2024-02
影响因子:
8.5
通讯作者:
Kai-Fung Chu;Weisi Guo
Kai-Fung Chu;Weisi Guo
中科院分区:
工程技术1区
文献类型:
--
作者:
Kai-Fung Chu;Weisi Guo

文献摘要

相似文献

移动即服务(MaaS)是一种新的传输模式,它将多种传输模式结合在一个平台上。基于过去经验的动态乘客行为需要对MaaS服务进行基于验证的优化。深度强化学习(DRL)可以通过基于个人乘客体验和偏好提供最合适的运输服务来提高乘客满意度。然而,这对使用集中式DRL方法的MaaS平台产生了新的隐私风险。如果平台没有精心设计隐私保护机制,就会发生信息泄露。在本文中,我们提出了一种联合深度确定性策略梯度(FDDPG),可以最大限度地提高乘客满意度和MaaS长期利润,同时保护隐私。我们实施了一个相等加权的经验采样机制,以防止采样偏差,FDDPG的解决方案质量是统计上等同于集中式算法。在模型训练和推理过程中,信息在本地进行处理,并且只共享梯度,这防止了信息泄漏给任何半诚实的参与者和窃听者。在梯度共享步骤中还采用了符合移动的Agent动态特性的安全聚合协议,以保证算法不受推理攻击。我们在纽约市的真实世界和合成场景上进行实验。实验结果表明,FDDPG算法可以使MaaS系统的利润和乘客满意度分别提高90%和15%,并能保持稳定的训练,防止Agent的流失。我们的方法和研究结果可以增强MaaS效用,并促进乘客对MaaS和其他数据驱动的交通系统的信任和参与。
Mobility-as-a-service (MaaS) is a new transport model that combines multiple transport modes in a single platform. Dynamic passenger behavior based on past experiences requires reinforcement-based optimization of MaaS services. Deep reinforcement learning (DRL) may improve passenger satisfaction by offering the most appropriate transport services based on individual passenger experiences and preferences. However, this produces a new privacy risk to the MaaS platform using the centralized DRL method. Information leakage will occur if the platform is not carefully designed with privacy-preserving mechanisms. In this paper, we propose a federated deep deterministic policy gradient (FDDPG) that maximizes passenger satisfaction and MaaS long-term profit while preserving privacy. We enforce an equally weighted experience sampling mechanism to prevent sampling bias such that the solution quality of FDDPG is statistically equivalent to the centralized algorithm. During the model training and inference, information is processed locally, and only the gradients are shared, which prevents information leakage to any semi-honest participants and eavesdroppers. Secure aggregation protocol in line with the dynamic property of the mobile agent is also used in the gradient sharing step to ensure that the algorithm is prevented from inference attacks. We perform experiments on New York City-based real-world and synthetic scenarios. The results show that the proposed FDDPG can improve the MaaS profit and passenger satisfaction by about 90% and 15%, respectively, and maintain stable training against agent dropout. Our approach and findings could enhance MaaS utility as well as facilitate passenger trust and participation in MaaS and other data-driven transportation systems.