Few-Shot Image-to-Semantics Translation for Policy Transfer in Reinforcement Learning

Few-Shot Image-to-Semantics Translation for Policy Transfer in Reinforcement Learning
复制标题

DOI:
10.1109/ijcnn55064.2022.9892464
复制
发表时间:
2022-07
期刊:
2022 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Reimi Sato;Kazuto Fukuchi;Jun Sakuma;Youhei Akimoto
Reimi Sato;Kazuto Fukuchi;Jun Sakuma;Youhei Akimoto
中科院分区:
其他
文献类型:
--
作者:
Reimi Sato;Kazuto Fukuchi;Jun Sakuma;Youhei Akimoto

文献摘要

相似文献

我们使用图像到语义翻译来研究策略转移,以减轻基于视觉的机器人控制代理的学习困难。该问题假设两个环境:以语义(即低维、本质信息)为状态空间的模拟器环境,以及以图像为状态空间的真实环境。通过学习从图像到语义的映射,我们可以将在模拟器中预先训练的策略转移到现实世界,从而消除现实世界中的策略代理交互学习,这是昂贵且有风险的。此外,与其他类型的模拟到真实的传输策略相比,使用图像到语义的映射在训练策略的计算效率和获得的策略的可解释性方面具有优势。为了解决学习图像到语义映射的主要困难,即生成训练数据集的人工注释成本,我们提出了两种技术:在模拟器环境中使用转换函数进行配对增强和主动学习。我们观察到注释成本降低了,而传输性能却没有下降,并且所提出的方法优于没有注释的现有方法。
We investigate policy transfer using image-to-semantics translation to mitigate learning difficulties in vision-based robotics control agents. This problem assumes two environments: a simulator environment with semantics, that is, low-dimensional and essential information, as the state space, and a real-world environment with images as the state space. By learning mapping from images to semantics, we can transfer a policy, pre-trained in the simulator, to the real world, thereby eliminating real-world on-policy agent interactions to learn, which are costly and risky. In addition, using image-to-semantics mapping is advantageous in terms of the computational efficiency to train the policy and the interpretability of the obtained policy over other types of sim-to-real transfer strategies. To tackle the main difficulty in learning image-to-semantics mapping, namely the human annotation cost for producing a training dataset, we propose two techniques: pair augmentation with the transition function in the simulator environment and active learning. We observed a reduction in the annotation cost without a decline in the performance of the transfer, and the proposed approach outperformed the existing approach without annotation.