Learning Q-network for Active Information Acquisition

Learning Q-network for Active Information Acquisition
复制标题

用于主动信息获取的学习 Q 网络

DOI:
--
复制
发表时间:
2019
期刊:
IEEE/RJS International Conference on Intelligent RObots and Systems
影响因子:
--
通讯作者:
George Pappas
George Pappas
中科院分区:
--
文献类型:
--
作者:
Heejin Jeong;Brent Schlotfeldt;Hamed Hassani;M. Morari;Daniel D. Lee;George Pappas

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新颖的加强学习方法,用于解决主动信息采集问题,该方法要求代理商选择一系列动作,以便使用板载传感器获取有关关注过程的信息。信息采集问题的经典挑战是计划算法对已知模型的依赖性以及计算信息理论成本功能在任意分布上的难度。相比之下,提议的增强学习框架不需要关于模型的任何知识,也不需要在扩展的训练阶段减轻问题。它导致政策有效地在线执行并适用于机器人系统的实时控制。此外,最先进的计划方法通常仅限于短范围,这可能会因本地最小值而有问题。强化学习自然会处理信息问题中规划范围的问题,因为它可以在长期有限或无限的时间范围内最大化折扣的奖励总和。我们讨论了提出的框架的潜在好处,并将新算法的性能与用于多目标跟踪方案的现有信息采集方法进行了比较。
In this paper, we propose a novel Reinforcement Learning approach for solving the Active Information Acquisition problem, which requires an agent to choose a sequence of actions in order to acquire information about a process of interest using on-board sensors. The classic challenges in the information acquisition problem are the dependence of a planning algorithm on known models and the difficulty of computing information-theoretic cost functions over arbitrary distributions. In contrast, the proposed framework of reinforcement learning does not require any knowledge on models and alleviates the problems during an extended training stage. It results in policies that are efficient to execute online and applicable for real-time control of robotic systems. Furthermore, the state-of-the-art planning methods are typically restricted to short horizons, which may become problematic with local minima. Reinforcement learning naturally handles the issue of planning horizon in information problems as it maximizes a discounted sum of rewards over a long finite or infinite time horizon. We discuss the potential benefits of the proposed framework and compare the performance of the novel algorithm to an existing information acquisition method for multi-target tracking scenarios.