Generating reward structures on a parameterized distribution of dynamics tasks

Generating reward structures on a parameterized distribution of dynamics tasks
复制标题

在动态任务的参数化分布上生成奖励结构

DOI:
10.1162/isal_a_00466
复制
发表时间:
2021
期刊:
Artificial Life Conference Proceedings
影响因子:
--
通讯作者:
Izquierdo, Eduardo J.
Izquierdo, Eduardo J.
中科院分区:
--
文献类型:
--
作者:
Leite, Abe;Izquierdo, Eduardo J.

文献摘要

参考文献

相似文献

为了使逼真的,多功能的学习适应人工领域,需要一组非常多样化的行为来学习。我们提出了一个经典的控制风格的任务与任务之间共享的信息最少的参数化分布。我们讨论是什么使一个任务琐碎,并提供一个基本的度量,收敛时间,衡量琐碎。然后,我们研究分析和经验的方法来生成奖励结构的任务的基础上,他们的动态,以尽量减少琐碎。与我们的预期相反,种群在奖励结构上进化,这种奖励结构激励状态空间中最稳定的位置花费最少的时间进行收敛,正如我们所定义的那样,因为我们的度量在这些背景下赋予行为微调的巨大重要性。这项工作铺平了道路,了解哪些任务分配,使学习的发展。
In order to make lifelike, versatile learning adaptive in the artificial domain, one needs a very diverse set of behaviors to learn. We propose a parameterized distribution of classic control-style tasks with minimal information shared between tasks. We discuss what makes a task trivial and offer a basic metric, time in convergence, that measures triviality. We then investigate analytic and empirical approaches to generating reward structures for tasks based on their dynamics in order to minimize triviality. Contrary to our expectations, populations evolved on reward structures that incentivized the most stable locations in state space spend the least time in convergence as we have defined it, because of the outsized importance our metric assigns to behavior fine-tuning in these contexts. This work paves the way towards an understanding of which task distributions enable the development of learning.
DOI: --
发表时间: 2019-01
期刊: ArXiv
影响因子: --
作者:
Rui Wang;J. Lehman;J. Clune;Kenneth O. Stanley
通讯作者: Rui Wang;J. Lehman;J. Clune;Kenneth O. Stanley
DOI: 10.1038/s41593-018-0310-2
发表时间: 2019-02-01
影响因子: 25
作者:
Yang, Guangyu Robert;Joglekar, Madhura R.;Wang, Xiao-Jing
通讯作者: Wang, Xiao-Jing
DOI: 10.1109/72.712153
发表时间: 1998-09-01
影响因子: --
作者:
Kodjabachian, J;Meyer, JA
通讯作者: Meyer, JA