The Sample Complexity of Teaching by Reinforcement on Q-Learning

The Sample Complexity of Teaching by Reinforcement on Q-Learning
复制标题

DOI:
10.1609/aaai.v35i12.17306
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Xuezhou Zhang;S. Bharti;Yuzhe Ma;A. Singla;Xiaojin Zhu
Xuezhou Zhang;S. Bharti;Yuzhe Ma;A. Singla;Xiaojin Zhu
中科院分区:
其他
文献类型:
--
作者:
Xuezhou Zhang;S. Bharti;Yuzhe Ma;A. Singla;Xiaojin Zhu

文献摘要

相似文献

我们研究了教学的样本复杂性,在文献中被称为‘教学维度’(TDim),用于强化教学范式,其中教师通过奖励来指导学生。这与由机器人应用程序推动的演示教学范式不同,在机器人应用程序中,教师通过提供状态/动作轨迹的演示进行教学。强化教学模式适用于更广泛的现实世界环境,在这些环境中,演示不方便,但没有得到系统的研究。本文重点研究了一类特殊的强化学习算法--Q-学习,刻画了不同教师对环境具有不同控制能力下的TDim,并给出了匹配的最优教学算法。我们的TDim结果提供了强化学习所需的最小样本数量,并讨论了它们与标准PAC风格的RL样本复杂性和示范样本复杂性结果的联系。我们的教学算法有可能在有帮助的老师的应用程序中加快RL代理的学习。
We study the sample complexity of teaching, termed as ``teaching dimension" (TDim) in the literature, for the teaching-by-reinforcement paradigm, where the teacher guides the student through rewards. This is distinct from the teaching-by-demonstration paradigm motivated by robotics applications, where the teacher teaches by providing demonstrations of state/action trajectories. The teaching-by-reinforcement paradigm applies to a wider range of real-world settings where a demonstration is inconvenient, but has not been studied systematically. In this paper, we focus on a specific family of reinforcement learning algorithms, Q-learning, and characterize the TDim under different teachers with varying control power over the environment, and present matching optimal teaching algorithms. Our TDim results provide the minimum number of samples needed for reinforcement learning, and we discuss their connections to standard PAC-style RL sample complexity and teaching-by-demonstration sample complexity results. Our teaching algorithms have the potential to speed up RL agent learning in applications where a helpful teacher is available.