Intelligent Systems and Pattern Recognition - Third International Conference, ISPR 2023, Hammamet, Tunisia, May 11-13, 2023, Revised Selected Papers, Part II

Intelligent Systems and Pattern Recognition - Third International Conference, ISPR 2023, Hammamet, Tunisia, May 11-13, 2023, Revised Selected Papers, Part II
复制标题

智能系统和模式识别 - 第三届国际会议,ISPR 2023,突尼斯哈马马特,2023 年 5 月 11-13 日,修订后的精选论文,第二部分

DOI:
10.1007/978-3-031-46338-9_12
复制
发表时间:
2024
期刊:
--
影响因子:
--
通讯作者:
Artaud C
Artaud C
中科院分区:
--
文献类型:
--
作者:
Artaud C

文献摘要

相似文献

人类的大脑赋予我们非凡的能力,使我们能够创造、想象和产生任何我们想要的东西。具体来说,我们拥有令人着迷的想象力,使我们能够从抽象概念中产生基本知识。在这些特征的激励下,机器学习的许多领域,尤其是无监督学习和强化学习,已经开始将这些想法作为核心。然而,这些方法并非没有缺点。强化学习的一个基本问题,特别是现在当神经网络作为函数逼近器使用时,是其可实现的最优性有限,相比于它在白板上的使用。由于神经网络学习的本质,每个任务可实现的行为是不一致的,提供一个统一的方法,使这种最优策略存在于参数空间中,将促进学习过程和行为结果。因此,我们感兴趣的是发现是否可以用无监督学习方法来促进强化学习,以减轻这种衰落。这项工作旨在分析使用生成模型提取学习到的强化学习策略(即模型参数)的可行性,目的是有条件地采样学习到的策略潜在空间以生成新策略。我们证明,在当前提出的体系结构下,这些模型能够在简单任务上重新创建策略,而在更复杂的任务上失败。因此,我们对这些失败进行了批判性分析,并讨论了进一步的改进,这将有助于这项工作的推广。
The human brain endows us with extraordinary capabilities that enable us to create, imagine, and generate anything we desire. Specifically, we have fascinating imaginative skills allowing us to generate fundamental knowledge from abstract concepts. Motivated by these traits, numerous areas of machine learning, notably unsupervised learning and reinforcement learning, have started using such ideas at their core. Nevertheless, these methods do not come without fault. A fundamental issue with reinforcement learning especially now when used with neural networks as function approximators is their limited achievable optimality compared to its uses from tabula rasa. Due to the nature of learning with neural networks, the behaviours achievable for each task are inconsistent and providing a unified approach that enables such optimal policies to exist within a parameter space would facilitate both the learning procedure and the behaviour outcomes. Consequently, we are interested in discovering whether reinforcement learning can be facilitated with unsupervised learning methods in a manner to alleviate this downfall. This work aims to provide an analysis of the feasibility of using generative models to extract learnt reinforcement learning policies (i.e. model parameters) with the intention of conditionally sampling the learnt policy-latent space to generate new policies. We demonstrate that under the current proposed architecture, these models are able to recreate policies on simple tasks whereas fail on more complex ones. We therefore provide a critical analysis of these failures and discuss further improvements which would aid the proliferation of this work.