A Partially Supervised Reinforcement Learning Framework for Visual Active Search

A Partially Supervised Reinforcement Learning Framework for Visual Active Search
复制标题

DOI:
10.48550/arxiv.2310.09689
复制
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Anindya Sarkar;Nathan Jacobs;Yevgeniy Vorobeychik
Anindya Sarkar;Nathan Jacobs;Yevgeniy Vorobeychik
中科院分区:
其他
文献类型:
--
作者:
Anindya Sarkar;Nathan Jacobs;Yevgeniy Vorobeychik

文献摘要

相似文献

视觉主动搜索(VAS)已被提出作为一个建模框架,其中视觉线索被用来指导探索,在一个大的地理空间区域的目标是确定感兴趣的区域。它的潜在应用包括识别珍稀野生动物偷猎活动的热点,搜索和救援场景,识别非法贩运武器,毒品或人口等。VAS的最新方法包括深度强化学习(DRL)的应用,它产生端到端的搜索策略,以及传统的主动搜索,它将预测与自定义算法方法相结合。虽然DRL框架已被证明在这些领域中大大优于传统的主动搜索,但其端到端的性质并没有充分利用在训练或实际搜索期间获得的监督信息,如果搜索任务与训练分布中的搜索任务显著不同,则这是一个显著的限制。我们提出了一种方法,结合了DRL和传统的主动搜索的强度分解成一个预测模块的搜索策略,它产生的地理空间分布的兴趣区域的任务嵌入和搜索历史的基础上,和一个搜索模块,它把预测和搜索历史作为输入,并输出搜索分布。我们开发了一种新的元学习方法,用于联合学习所产生的组合策略,该策略可以有效地利用在训练和决策时获得的监督信息。我们广泛的实验表明,所提出的表示和元学习框架显着优于最先进的视觉主动搜索在几个问题领域。
Visual active search (VAS) has been proposed as a modeling framework in which visual cues are used to guide exploration, with the goal of identifying regions of interest in a large geospatial area. Its potential applications include identifying hot spots of rare wildlife poaching activity, search-and-rescue scenarios, identifying illegal trafficking of weapons, drugs, or people, and many others. State of the art approaches to VAS include applications of deep reinforcement learning (DRL), which yield end-to-end search policies, and traditional active search, which combines predictions with custom algorithmic approaches. While the DRL framework has been shown to greatly outperform traditional active search in such domains, its end-to-end nature does not make full use of supervised information attained either during training, or during actual search, a significant limitation if search tasks differ significantly from those in the training distribution. We propose an approach that combines the strength of both DRL and conventional active search by decomposing the search policy into a prediction module, which produces a geospatial distribution of regions of interest based on task embedding and search history, and a search module, which takes the predictions and search history as input and outputs the search distribution. We develop a novel meta-learning approach for jointly learning the resulting combined policy that can make effective use of supervised information obtained both at training and decision time. Our extensive experiments demonstrate that the proposed representation and meta-learning frameworks significantly outperform state of the art in visual active search on several problem domains.