课题基金 / 基金详情

Flexible State Representations in Reinforcement Learning

Flexible State Representations in Reinforcement Learning
强化学习中灵活的状态表示
批准号:
0413004
负责人:
Satinder Baveja
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-02-15 至 2010-01-31

项目摘要

项目成果

Satinder Baveja的其他基金

相似基金

相关文献

中文摘要
翻译
强化学习(RL)适用于任何涉及代理采取一系列动作的任务,其中一个动作的效果会影响后续动作的长期效用。这样的学习代理应该如何表示它对环境的知识?传统的模型通常捕获代理的状态,包括环境中的对象和事件以及它们之间的关系。由于这些关系不能被代理人通过其传感器直接观察到,它们只在代理人的人类设计者的头脑中有意义。最近,PI和同事们提出将代理的状态建模为由代理可以在其环境中执行的测试或实验的可观察结果的一组预测组成。这种表示被称为预测状态表示(PSR),完全由可观察的量组成,并且在RL任务中有效和可扩展的规划和学习方面有很大的希望。在这个项目中,许多关于PSR的基本问题正在探索。这些问题包括:(1)智能体如何发现它应该保持什么样的预测来捕捉环境的状态?(2)作为PSR状态表示的一部分的长期行动条件预测如何用于加速评估行动长期影响的规划?(3)如何将过去观测的记忆与未来观测的PSR预测结合起来,以获得计算上的好处?以及(4)如何将灵活的状态时间抽象表示(特别是PSR)与类似的动作时间抽象表示相结合?该项目正在将PSRs的新生概念发展成为一个成熟的学习和规划理论。如果成功,这项研究将导致RL在人工智能,运筹学,控制和动态系统的大规模领域中构建学习代理的适用性大幅增加。该项目还计划构建一组基准RL任务并使其可用。这将有助于弥补RL社区缺乏这种广泛可用的测试床的问题。
英文摘要
Reinforcement learning (RL) applies to any task that involves an agent taking a sequence of actions where the effects of one action influence the long-term utility of subsequent actions. How should such a learning agent represent its knowledge about its environment? Traditional models typically capture the agent's state as composed of objects and events in the environment and relations among them. Since these relations cannot be directly observed by the agent through its sensors, they have meaning only in the mind of the human designer of the agent. Recently, the PI and colleagues have instead proposed modeling the agent's state as composed of a set of predictions of observable outcomes of tests or experiments that the agent could perform in its environment. Such representations, called predictive state representations (PSRs), are composed entirely of observable quantities and therein lies much of their promise for efficient and scalable planning and learning in RL tasks. In this project many foundational questions about PSRs are being explored. These include: (1) How can an agent discover what predictions it should keep to capture the state of its environment?; (2) How can the long-term action-conditional predictions that are part of PSR state representations be used to speed up planning that is about evaluating long-term effects of actions?; (3) How can memory of past observations be combined with PSR predictions of future observations for computational benefit?; and (4) How can the flexible temporally abstract representations of state -- specifically PSRs -- be combined with similarly temporally abstract representations of actions? This project is developing the nascent idea of PSRs into a full-fledged theory of learning and planning. If successful, this research will result in a dramatic increase in the applicability of RL for building learning agents in large-scale domains in AI, operations research, control, and dynamical systems. This project also plans to construct and make publically available a set of benchmark RL tasks. This will help remediate the lack of such widely available test beds in the RL community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Combining Reinforcement Learning and Deep Learning Methods to Address High-Dimensional Perception, Partial Observability and Delayed Reward
RI: Small: Reinforcement Learning with Predictive State Representations
EAGER: On the Optimal Rewards Problem
SHB: Medium: Collaborative Research: Novel Computational Techniques for Cardiovascular Risk Stratification
国内基金
海外基金
Simulation and certification of the ground state of many-body systems on quantum simulators
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Abolfazl Bayat
  • 依托单位:
Cortical control of internal state in the insular cortex-claustrum region
微波有源Scattering dark state粒子的理论及应用研究
  • 批准号:
    61701437
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    28.0万元
  • 批准年份:
    2017
  • 负责人:
    李欢
  • 依托单位: