Hierarchical State Abstraction Based on Structural Information Principles

Hierarchical State Abstraction Based on Structural Information Principles
复制标题

DOI:
10.48550/arxiv.2304.12000
复制
发表时间:
2023-04
期刊:
--
影响因子:
--
通讯作者:
Xianghua Zeng;Hao Peng;Angsheng Li;Chunyang Liu;Lifang He;Philip S. Yu
Xianghua Zeng;Hao Peng;Angsheng Li;Chunyang Liu;Lifang He;Philip S. Yu
中科院分区:
其他
文献类型:
--
作者:
Xianghua Zeng;Hao Peng;Angsheng Li;Chunyang Liu;Lifang He;Philip S. Yu

文献摘要

相似文献

状态抽象通过在具有丰富观察的强化学习中忽略不相关的环境信息来优化决策。然而,最近的方法侧重于足够的代表性能力,导致重要的信息丢失,影响他们在挑战性任务中的表现。在本文中,我们从信息论的角度提出了一种新颖的基于数学结构信息原理的状态抽象框架,即SISA。具体来说,提出了一种无需人工辅助的无监督自适应分层状态聚类方法,同时生成最优编码树。在每个非根树节点上,设计了新的聚合函数和条件结构熵,以实现分层状态抽象并补偿状态抽象中采样引起的基本信息损失。对视觉网格世界域和六个连续控制基准的实证评估表明,与五种 SOTA 状态抽象方法相比,SISA 显着提高了平均事件奖励和样本效率,分别高达 18.98 和 44.44%。此外,我们通过实验证明 SISA 是一个通用框架,可以灵活地与不同的表示学习目标集成,以进一步提高其性能。
State abstraction optimizes decision-making by ignoring irrelevant environmental information in reinforcement learning with rich observations. Nevertheless, recent approaches focus on adequate representational capacities resulting in essential information loss, affecting their performances on challenging tasks. In this article, we propose a novel mathematical Structural Information principles-based State Abstraction framework, namely SISA, from the information-theoretic perspective. Specifically, an unsupervised, adaptive hierarchical state clustering method without requiring manual assistance is presented, and meanwhile, an optimal encoding tree is generated. On each non-root tree node, a new aggregation function and condition structural entropy are designed to achieve hierarchical state abstraction and compensate for sampling-induced essential information loss in state abstraction. Empirical evaluations on a visual gridworld domain and six continuous control benchmarks demonstrate that, compared with five SOTA state abstraction approaches, SISA significantly improves mean episode reward and sample efficiency up to 18.98 and 44.44%, respectively. Besides, we experimentally show that SISA is a general framework that can be flexibly integrated with different representation-learning objectives to improve their performances further.