Co-training for Policy Learning

Co-training for Policy Learning
复制标题

DOI:
--
复制
发表时间:
2019-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Jialin Song;Ravi Lanka;Yisong Yue;M. Ono
Jialin Song;Ravi Lanka;Yisong Yue;M. Ono
中科院分区:
其他
文献类型:
--
作者:
Jialin Song;Ravi Lanka;Yisong Yue;M. Ono

文献摘要

相似文献

我们研究了在多个状态-动作表示的环境中学习顺序决策策略的问题。这样的设置自然出现在许多领域,例如规划(例如,多个整数规划公式)和各种组合优化问题(例如,具有整数编程和基于图形的公式化的那些)。受经典分类协同训练框架的启发,本文研究了策略学习协同训练问题。我们提出了充分的条件下,学习从两个视图可以提高学习从一个单一的视图。受这些理论见解的启发,我们提出了一个元算法,用于顺序决策的协同训练。我们的框架兼容强化学习和模仿学习。我们验证了我们的方法在广泛的任务,包括离散/连续控制和组合优化的有效性。
We study the problem of learning sequential decision-making policies in settings with multiple state-action representations. Such settings naturally arise in many domains, such as planning (e.g., multiple integer programming formulations) and various combinatorial optimization problems (e.g., those with both integer programming and graph-based formulations). Inspired by the classical co-training framework for classification, we study the problem of co-training for policy learning. We present sufficient conditions under which learning from two views can improve upon learning from a single view alone. Motivated by these theoretical insights, we present a meta-algorithm for co-training for sequential decision making. Our framework is compatible with both reinforcement learning and imitation learning. We validate the effectiveness of our approach across a wide range of tasks, including discrete/continuous control and combinatorial optimization.