Learning an Urban Air Mobility Encounter Model from Expert Preferences

Learning an Urban Air Mobility Encounter Model from Expert Preferences
复制标题

根据专家偏好学习城市空中交通遭遇模型

DOI:
10.1109/dasc43569.2019.9081648
复制
发表时间:
2019
期刊:
2019 IEEE/AIAA 38th Digital Avionics Systems Conference (DASC)
影响因子:
--
通讯作者:
Mykel J. Kochenderfer
Mykel J. Kochenderfer
中科院分区:
--
文献类型:
--
作者:
Sydney M. Katz;Anne;Mykel J. Kochenderfer

文献摘要

被引文献

相似文献

空域模型在有人和无人驾驶飞机的飞机防撞系统的开发和评估中发挥了重要作用。随着城市空中机动性(UAM)系统的发展,我们需要能够代表其作战环境的新的相遇模型。由于缺乏关于UAM在空域中行为的数据,开发此类模型具有挑战性。以前针对其他机型的相遇模型依赖于大型数据集来生成逼真的轨迹,而本文提出了一种基于专家知识的相遇建模方法。特别是,基于偏好的学习的最新进展被扩展到根据专家偏好调整相遇模型。该模型采用马尔可夫决策过程(MDP)的随机策略的形式,其中报酬函数是从领域专家的两两查询中学习的。我们评估了两种查询方法的性能,这两种方法寻求最大化从每个查询获得的信息。最终,我们演示了一种只需专家几分钟的时间就可以生成真实相遇轨迹的方法。
Airspace models have played an important role in the development and evaluation of aircraft collision avoidance systems for both manned and unmanned aircraft. As Urban Air Mobility (UAM) systems are being developed, we need new encounter models that are representative of their operational environment. Developing such models is challenging due to the lack of data on UAM behavior in the airspace. While previous encounter models for other aircraft types rely on large datasets to produce realistic trajectories, this paper presents an approach to encounter modeling that instead relies on expert knowledge. In particular, recent advances in preference-based learning are extended to tune an encounter model from expert preferences. The model takes the form of a stochastic policy for a Markov decision process (MDP) in which the reward function is learned from pairwise queries of a domain expert. We evaluate the performance of two querying methods that seek to maximize the information obtained from each query. Ultimately, we demonstrate a method for generating realistic encounter trajectories with only a few minutes of an expert's time.