课题基金 / 基金详情

Robust Policy Learning

Robust Policy Learning
稳健的政策学习
批准号:
2579024
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
关键词:

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
简要描述研究背景,包括潜在影响我对从大规模离线数据集学习时寻找稳健和因果正确策略的问题感兴趣。各种机器人数据集中场景的倾斜分布意味着存在一种不平衡,大多数有趣的场景都在大尾巴上,而数据集中的大多数样本涵盖了一些默认的琐碎行为(例如,在驾驶任务中“直线巡航”,而尾巴由转弯、刹车等组成)。这种不平衡导致了一系列的病态:从样本低效的训练到从数据中学习虚假的相关性。这些构成了一系列有趣的研究问题,需要逐步解决。这将对机器人学习和控制领域产生很大的影响,因为与环境或专家演示者的交互是昂贵的,样本效率将帮助这些学习算法扩展到以前无法进入的领域,比如那些对安全至关重要的领域。目的和目标:发表我在以下领域的研究:从示范中学习,持续学习,因果正确的政策学习。协同工作,提出有意义的基准,以促进这些领域的进步。研究方法的新颖性:我计划使用因果推理和贝叶斯深度学习框架的方法来解决当前训练范式的局限性,并解决上述问题。与EPSRC的战略和研究领域(项目涉及的EPSRC研究领域)保持一致:工程领域的进一步信息可在http://www.epsrc.ac.uk/research/ourportfolio/researchareas/Any上找到,涉及的公司或合作者:丰田欧洲和丰田研究,美国
英文摘要
Brief description of the context of the research including potential impactI'm interested in the problem of finding robust and causally-correct policies when learning from large-scale offline datasets. The skewed distribution of scenarios in various robotics datasets means that there is an imbalance where most of the interesting scenarios are in the heavy tail whereas the majority of samples in the dataset cover some default trivial behavior (for example, 'cruising straight' in the driving task, while the tail consists of turns, braking and so on). This imbalance results in a range of pathologies: from sample inefficient training to learning spurious correlations from the data. These make for an interesting set of research problems to tackle step by step. This would have a lot of impact in the field of robot learning and control, since interaction with an environment or expert demonstrator is expensive, and sample efficiency will help these learning algorithms scale to previously inaccessible domains like those that are safety critical.Aims and Objectives : Publish my research in the following areas: learning from demonstrations, continual learning, causally-correct policy learningWork collaboratively to propose meaningful benchmarks to spur progress in these areasNovelty of the research methodology : I plan to employ methodologies from the causal inference and bayesian deep learning frameworks to address the limitations of the current training paradigm and solve the aforementioned problems. Alignment to EPSRC's strategies and research areas (which EPSRC research area the project relates to) : engineeringFurther information on the areas can be found on http://www.epsrc.ac.uk/research/ourportfolio/researchareas/Any companies or collaborators involved: Toyota Europe and Toyota Research, USA
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
The Heterogenous Impact of Monetary Policy on Firms' Risk and Fundamentals
Financial Constraints in China and Their Policy Implications
  • 批准号:
    --
  • 项目类别:
    外国优秀青年学 者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    Jake Zhao
  • 依托单位: