课题基金 / 基金详情

Bayesian Deep Reinforcement Learning

Bayesian Deep Reinforcement Learning
贝叶斯深度强化学习
批准号:
2243850
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在机器人控制、围棋和自动驾驶等具有挑战性的环境中,深度强化学习在学习控制策略方面已经无处不在。然而,标准方法通常是低效的、不稳定的,并且使用的特殊技巧在理论上是不合理的。该项目着眼于通过贝叶斯透镜推导深度强化学习的原则性新目标和算法,以及对现有强化学习算法的新解释。目的和目标-通过贝叶斯自适应MDP框架中的元学习将贝叶斯强化学习中的现有方法扩展到更复杂的环境。-研究深度强化学习和概率推理之间的联系,并以此得出强化学习的原则性目标。我们的目标是为现有的目标带来理论证明,如均方Bellman误差,它可能不能适当地反映函数空间的几何形状。-扩展相关方法,如贝叶斯优化到新的设置。研究方法的新颖性-强化学习的概率处理-新颖和新提出的强化学习框架,如bamdp和RL作为推论与EPSRC的战略和研究领域(EPSRC研究领域与项目相关)保持一致,有关该领域的进一步信息可在http://www.epsrc.ac.uk/research/ourportfolio/researchareas/-上找到人工智能技术-机器人-统计和应用概率任何公司或合作者involvedNone
英文摘要
Brief description of the context of the research including potential impactDeep reinforcement learning has become ubiquitous for learning control policies in challenging environments such as robotic control, Go-playing and autonomous driving. However, standard approaches are often sample inefficient, unstable, and use ad-hoc tricks that are not theoretically well justified. This project looks at deriving principled new objectives and algorithms for deep reinforcement learning through a Bayesian lens, and new interpretations of existing reinforcement learning algorithms.Aims and Objectives- Extending existing methods in the Bayesian Reinforcement Learning via meta-learning in the Bayes Adaptive MDP Framework to more complex environments.- Studying the connection between deep reinforcement learning and probabilistic inference and using this to derive a principled objective for reinforcement learning. We aim to bring theoretical justification for existing objectives such as mean-squared Bellman error which may not appropriately reflect the geometry of the function space.- Scaling related methodologies such as Bayesian optimisation to novel settings.Novelty of the research methodology- Probabilistic treatment of reinforcement learning- Novel and newly proposed frameworks for reinforcement learning such as BAMDPs and RL as inferenceAlignment to EPSRC's strategies and research areas (which EPSRC research area the project relates to) Further information on the areas can be found on http://www.epsrc.ac.uk/research/ourportfolio/researchareas/- Artificial intelligence technologies- Robotics- Statistics and applied probabilityAny companies or collaborators involvedNone
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
  • 批准号:
    2026JJ81909
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    胡曦
  • 依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
  • 批准号:
    12271434
  • 项目类别:
    面上项目
  • 资助金额:
    46万元
  • 批准年份:
    2022
  • 负责人:
    贺小伟
  • 依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
  • 批准号:
    2020A151501709
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2020
  • 负责人:
    谢怡
  • 依托单位:
面向Deep Web的数据整合关键技术研究
  • 批准号:
    61872168
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2018
  • 负责人:
    董永权
  • 依托单位: