课题基金 / 基金详情

Sequential Decision making in probabilistic models

Sequential Decision making in probabilistic models
概率模型中的顺序决策
批准号:
2744311
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
已结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该建议考虑了非线性环境下的鲁棒序贯决策问题。强化学习在解决非线性环境中的复杂问题方面表现出很高的潜力,但缺乏效率和鲁棒性。我们认为,为了在真实的世界中部署强化学习代理,必须开发类似的效率和鲁棒性,已经在控制理论中开发。我们提出了广泛的控制和概率推理文献,以改善RL算法,并提出了两个有趣的研究方向。第一个考虑使用顺序蒙特-卡罗方法来改善非线性规划。第二个方向侧重于通过探索对抗学习,鲁棒控制理论和不确定性建模之间的联系来设计鲁棒控制器。
英文摘要
This proposal considers the problem of robust sequential decision making in non-linear environments. Reinforcementlearning has demonstrated high potential for solving complex problems in non-linear environments but has lackedefficiency and robustness. We argue that in order to deploy reinforcement learning agents in the real world, it is essential todevelop similar efficiency and robustness properties that have been developed in control theory. We propose to leveragethe extensive control and probabilistic reasoning literature to improve RL algorithms and present two interesting researchdirections. The first one considers using Sequential Monte-Carlo methods to improve planning for non-linearenvironments. The second direction focuses on designing robust controllers by exploring the connections betweenadversarial learning, robust control theory, and uncertainty modelling.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis