Constructive Policy: Reinforcement Learning Approach for Connected Multi-Agent Systems

Constructive Policy: Reinforcement Learning Approach for Connected Multi-Agent Systems
复制标题

DOI:
10.1109/coase.2019.8843223
复制
发表时间:
2019-08
期刊:
2019 IEEE 15th International Conference on Automation Science and Engineering (CASE)
影响因子:
--
通讯作者:
Sayyed Jaffar Ali Raza;Mingjie Lin
Sayyed Jaffar Ali Raza;Mingjie Lin
中科院分区:
其他
文献类型:
--
作者:
Sayyed Jaffar Ali Raza;Mingjie Lin

文献摘要

被引文献

相似文献

基于策略的增强学习方法被广泛用于以部分或没有模型表示的多种状态来学习最佳动作。受到启发(A)蛇或蛇形机器人)在此类环境中表现出有限的性能,这是由于环境的稀疏性质,并且在本文中没有完全可观察的模型表示。通过将其分解为相同的,连接和缩放的多基因结构,然后将学习框架应用于本地和全球排名的层次。本地层的偏见函数涉及“可重复使用”本地代理的“可重复使用”,以实现本地策略,也可以由其他相同的本地代理人重复使用。一项政策,应用当地代理的整个连接结构的当地政策的正确组合,以通过学习本地策略并学习全球政策后,通过协作构建全球任务。代理商,基于最佳子任务的最大奖励,将当地代理人作为对全球奖励的积极偏见具有超冗余自由度(DOF)的蛇形机器人的建设性策略方法,用于实现最佳控制,我们还概述了与分层学徒学习方法的连接,可以将其视为复杂控制任务的分层学习框架。
Policy based reinforcement learning methods are widely used for multi-agent systems to learn optimal actions given any state; with partial or even no model representation. However multi-agent systems with complex structures (curse of dimensionality) or with high constraints (like bio-inspired (a) snake or serpentine robots) show limited performance in such environments due to sparse-reward nature of environment and no fully observable model representation. In this paper we present a constructive learning and planning scheme that reduces the complexity of high-diemensional agent model by decomposing it into identical, connected and scaled down multiagent structure and then apply learning framework in layers of local and global ranking. Our layered hierarchy method also decomposes the final goal into multiple sub-tasks and a global task (final goal) that is bias-induced function of local sub-tasks. Local layer deals with learning ‘reusable’ local policy for a local agent to achieve a sub-task optimally; that local policy can also be reused by other identical local agents. Furthermore, global layer learns a policy to apply right combination of local policies that are parameterized over entire connected structure of local agents to achieve the global task by collaborative construction of local agents. After learning local policies and while learning global policy, the framework generates sub-tasks for each local agent, and accepts local agents’ intrinsic rewards as positive bias towards maximum global reward based of optimal sub-tasks assignments. The advantage of proposed approach includes better exploration due to decomposition of dimensions, and reusability of learning paradigm over extended dimension spaces. We apply the constructive policy method to serpentine robot with hyper-redundant degrees of freedom (DOF), for achieving optimal control and we also outline connection to hierarchical apprenticeship learning methods which can be seen as layered learning framework for complex control tasks.