TrajGAIL: Trajectory Generative Adversarial Imitation Learning for Long-term Decision Analysis

TrajGAIL: Trajectory Generative Adversarial Imitation Learning for Long-term Decision Analysis
复制标题

DOI:
10.1109/icdm50108.2020.00089
复制
发表时间:
2020-11
期刊:
2020 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Xin Zhang;Yanhua Li;Xun Zhou;Ziming Zhang;Jun Luo
Xin Zhang;Yanhua Li;Xun Zhou;Ziming Zhang;Jun Luo
中科院分区:
其他
文献类型:
--
作者:
Xin Zhang;Yanhua Li;Xun Zhou;Ziming Zhang;Jun Luo

文献摘要

相似文献

移动的传感和信息技术使我们能够从人类决策者那里收集大量的移动数据,例如,出租车、Uber汽车的GPS轨迹,以及乘坐公共汽车和火车的乘客行程数据。从这些数据中理解和学习人类的决策策略,可以潜在地促进个人的福祉,提高运输服务质量。现有的人类策略学习工作,如反向强化学习,都将决策过程建模为马尔可夫决策过程,从而假设马尔可夫属性。在这项工作中,我们表明,这样的马尔可夫属性不持有在现实世界中的人类决策过程。为了应对这一挑战,我们开发了一个轨迹生成对抗模仿学习(TrajGAIL)框架。它通过将人类决策过程建模为可变长度马尔可夫决策过程(VLMDP)来捕获长期决策依赖性,并设计了一个基于深度神经网络的框架,以从人类代理的历史数据集逆向学习决策策略。我们验证我们的框架使用两个真实的世界人类生成的时空数据集,包括出租车司机寻找司机的决策数据和公共交通出行数据。结果表明,与马尔可夫属性假设的基线相比,在学习人类决策策略方面有显着的准确性提高。
Mobile sensing and information technology have enabled us to collect a large amount of mobility data from human decision-makers, for example, GPS trajectories from taxis, Uber cars, and passenger trip data of taking buses and trains. Understanding and learning human decision-making strategies from such data can potentially promote individual's well-being and improve the transportation service quality. Existing works on human strategy learning, such as inverse reinforcement learning, all model the decision-making process as a Markov decision process, thus assuming the Markov property. In this work, we show that such Markov property does not hold in real-world human decision-making processes. To tackle this challenge, we develop a Trajectory Generative Adversarial Imitation Learning (TrajGAIL) framework. It captures the long-term decision dependency by modeling the human decision processes as variable length Markov decision processes (VLMDPs), and designs a deep-neural-network-based framework to inversely learn the decision-making strategy from the human agent's historical dataset. We validate our framework using two real world human-generated spatial-temporal datasets including taxi driver passenger-seeking decision data and public transit trip data. Results demonstrate significant accuracy improvement in learning human decision-making strategies, when comparing to baselines with Markov property assumptions.