Learning Submodular Objectives for Team Environmental Monitoring

Learning Submodular Objectives for Team Environmental Monitoring
复制标题

学习团队环境监测的子模块目标

DOI:
--
复制
发表时间:
2021
影响因子:
5.2
通讯作者:
Stephen L. Smith
Stephen L. Smith
中科院分区:
计算机科学2区
文献类型:
--
作者:
Nils Wilde;Armin Sadeghi;Stephen L. Smith

文献摘要

参考文献

被引文献

相似文献

在本文中,我们研究了著名的团队定向运动的问题,其中一队机器人收集奖励访问的位置。通常情况下,奖励被认为是已知的机器人;然而,在应用程序,如环境监测或场景重建,奖励往往是主观的,并指定它们是具有挑战性的。我们提出了一个框架来学习用户的未知偏好,通过向他们提供替代解决方案,用户提供了一个建议的替代解决方案的排名。我们考虑两种情况下的用户:1)确定性用户提供的替代解决方案的最佳排名,和2)噪声用户提供的最佳排名根据一个未知的概率分布。对于确定性的用户,我们提出了一个框架,以尽量减少从最佳解决方案,即遗憾的最大偏差的约束。我们调整的方法来捕捉嘈杂的用户,并尽量减少预期的遗憾。最后,我们证明了学习用户偏好的重要性和所提出的方法的性能在一组广泛的实验结果,使用真实的世界数据集环境监测问题。
In this paper, we study the well-known team orienteering problem where a fleet of robots collects rewards by visiting locations. Usually, the rewards are assumed to be known to the robots; however, in applications such as environmental monitoring or scene reconstruction, the rewards are often subjective and specifying them is challenging. We propose a framework to learn the unknown preferences of the user by presenting alternative solutions to them, and the user provides a ranking on the proposed alternative solutions. We consider the two cases for the user: 1) a deterministic user which provides the optimal ranking for the alternative solutions, and 2) a noisy user which provides the optimal ranking according to an unknown probability distribution. For the deterministic user we propose a framework to minimize a bound on the maximum deviation from the optimal solution, namely regret. We adapt the approach to capture the noisy user and minimize the expected regret. Finally, we demonstrate the importance of learning user preferences and the performance of the proposed methods in an extensive set of experimental results using real world datasets for environmental monitoring problems.
DOI: 10.1109/icra.2018.8460854
发表时间: 2018-05
期刊: 2018 IEEE International Conference on Robotics and Automation (ICRA)
影响因子: --
作者:
Yuchen Cui;S. Niekum
通讯作者: Yuchen Cui;S. Niekum
提出简单的问题:一种用户友好的主动奖励学习方法
DOI: --
发表时间: 2019
期刊: Proceedings of the 3rd Conference on Robot Learning
影响因子: --
作者:
Erdem Biyik, Malayandi Palan
通讯作者: Erdem Biyik, Malayandi Palan