Generating reward structures on a parameterized distribution of dynamics tasks
Generating reward structures on a parameterized distribution of dynamics tasks
复制标题
在动态任务的参数化分布上生成奖励结构
DOI:
10.1162/isal_a_00466
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Izquierdo, Eduardo J.
中科院分区:
文献类型:
--
作者:
Leite, Abe;Izquierdo, Eduardo J.
In order to make lifelike, versatile learning adaptive in the artificial domain, one needs a very diverse set of behaviors to learn. We propose a parameterized distribution of classic control-style tasks with minimal information shared between tasks. We discuss what makes a task trivial and offer a basic metric, time in convergence, that measures triviality. We then investigate analytic and empirical approaches to generating reward structures for tasks based on their dynamics in order to minimize triviality. Contrary to our expectations, populations evolved on reward structures that incentivized the most stable locations in state space spend the least time in convergence as we have defined it, because of the outsized importance our metric assigns to behavior fine-tuning in these contexts. This work paves the way towards an understanding of which task distributions enable the development of learning.
DOI:
--
发表时间:
2019-01
期刊:
ArXiv
影响因子:
--
作者:
Rui Wang;J. Lehman;J. Clune;Kenneth O. Stanley
通讯作者:
Rui Wang;J. Lehman;J. Clune;Kenneth O. Stanley
影响因子:
25
作者:
Yang, Guangyu Robert;Joglekar, Madhura R.;Wang, Xiao-Jing
通讯作者:
Wang, Xiao-Jing
影响因子:
--
作者:
Kodjabachian, J;Meyer, JA
通讯作者:
Meyer, JA