课题基金 / 基金详情

RI: Small: Feature Encoding for Reinforcement Learning

RI: Small: Feature Encoding for Reinforcement Learning
RI:小型:强化学习的特征编码
批准号:
1815300
负责人:
Ronald Parr
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-01 至 2023-07-31

项目摘要

项目成果

Ronald Parr的其他基金

相似基金

相关文献

中文摘要
翻译
该项目专注于机器学习的子领域,称为强化学习(RL),其中算法或机器人通过试错来学习。与机器学习的许多领域一样,人们对强化学习的“深度学习”方法(即“深度RL”)的兴趣激增。“深度学习使用由动物大脑中发现的结构驱动的计算模型。Deep RL已经取得了一些惊人的成功,包括最近的一项进展,一个程序学会了比最好的人类玩家更好地玩亚洲围棋。值得注意的是,这种性能水平是在没有任何人为指导的情况下实现的。只给游戏规则,程序通过与自己博弈来学习。虽然游戏很有趣,也很吸引人,但这只是一个技术演示。企业正在寻求部署Deep RL方法,以提高其在数据中心管理和机器人等一系列应用中的运营效率。为了充分发挥深度强化学习的潜力,需要进一步的研究,使训练过程更加可预测、可靠和高效。目前的技术需要大量的训练数据和计算,系统配置的细微变化可能会导致所获得的结果质量的巨大差异。因此,即使RL系统可以通过试错自主学习,但可能需要大量的人类直觉,经验和实验才能为这些系统的成功奠定基础。该提案旨在开发新的技术和理论,使高质量的深度RL结果更广泛,更容易获得。此外,该提案将为本科生提供机会,通过杜克的数据+倡议参与研究。拟议的研究部分受到过去关于强化学习的特征选择和发现的工作的启发。 大部分的工作主要集中在线性值函数近似。 它与深度强化学习的相关性在于,诸如深度Q学习之类的方法具有线性最终层。 因此,前面的非线性层可以被解释为执行特征发现,最终是线性值函数逼近过程。 在早期工作中为成功的线性值函数近似指定的特征的充分条件现在可以重新解释为深度网络倒数第二层的中间目标函数。拟议的研究旨在实现以下目标:1)开发一种解释和通知深度强化学习方法的特征构建理论,2)开发适用于深度强化学习的改进的值函数近似方法,3)开发适用于深度强化学习的改进的策略搜索方法,以及4)开发用于强化学习中的探索的新算法,其利用学习的特征表示,和5)进行计算实验,证明新算法在基准问题上的有效性。该奖项反映了NSF的法定使命,并已被视为通过使用基金会的知识价值和更广泛的影响审查标准进行评估,
英文摘要
This project focuses on the subfield of machine learning referred to as Reinforcement Learning (RL), in which algorithms or robots learn by trial and error. As with many areas of machine learning, there has been a surge of interest in "deep learning" approaches to reinforcement learning, i.e, "Deep RL." Deep learning uses computational models motivated by structures found in the brains of animals. Deep RL has enjoyed some stunning successes, including a recent advance by which a program learned to play the Asian game of Go better than the best human player. Notably, this level of performance was achieved without any human guidance. Given only the rules of the game, the program learned by playing against itself. Although games are intriguing and attention-grabbing, this feat was merely a technology demonstration. Firms are seeking to deploy Deep RL methods to increase the efficiency of their operations across a range of applications such as data center management and robotics. To realize fully the potential of Deep RL, further research is required to make the training process more predictable, reliable, and efficient. Current techniques require massive amounts of training data and computation, and subtle changes in the configuration of the system can cause huge differences in the quality of the results obtained. Thus, even though RL systems can learn autonomously by trial and error, a large amount of human intuition, experience and experimentation may be required to lay the groundwork for these systems to succeed. This proposal seeks to develop new techniques and theory to make high quality deep RL results more widely and easily obtainable. In addition, this proposal will provide opportunities for undergraduates to be involved in research through Duke's Data+ initiative.The proposed research is partly inspired by past work on feature selection and discovery for reinforcement learning. Much of that work focused primarily on linear value function approximation. Its relevance to deep reinforcement learning is that methods such as Deep Q-learning have a linear final layer. The preceding, nonlinear layers can, therefore, be interpreted as performing feature discovery for what is ultimately a linear value function approximation process. Sufficient conditions on the features that were specified for successful linear value function approximation in earlier work can now be re-interpreted as an intermediate objective function for the penultimate layer of a deep network. The proposed research aims to achieve the following objectives: 1) Develop a theory of feature construction that explains and informs deep reinforcement learning methods, 2) develop improved approaches to value function approximation that are applicable to deep reinforcement learning, 3) develop improved approaches to policy search that are applicable to deep reinforcement learning, and 4) develop new algorithms for exploration in reinforcement learning that take advantage of learned feature representations, and 5) perform computational experiments demonstrating the efficacy of the new algorithms developed on benchmark problems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2021
期刊:
影响因子: --
作者: [Mark W. Nemecek;R. Parr]
通讯作者: Mark W. Nemecek;R. Parr
EAGER: Collaborative Research: An Unified Learnable Roadmap for Sequential Decision Making in Relational Domains
  • 批准号:
    1836575
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2018
  • 负责人:
    Ronald Parr
  • 依托单位:
RI: Small: Non-parametric Approximate Dynamic Programming for Continuous Domains
  • 批准号:
    1218931
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2012
  • 负责人:
    Ronald Parr
  • 依托单位:
EAGER: IIS: RI: Learning in Continuous and High Dimensional Action Spaces
  • 批准号:
    1147641
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2011
  • 负责人:
    Ronald Parr
  • 依托单位:
Collaborative: RI: Feature Discovery and Benchmarks for Exportable Reinforcement Learning
  • 批准号:
    0713435
  • 项目类别:
    Standard Grant
  • 资助金额:
    $22.5万
  • 财政年份:
    2007
  • 负责人:
    Ronald Parr
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: