课题基金 / 基金详情

Object-Centric Visual Representation And Reinforcement Learning

Object-Centric Visual Representation And Reinforcement Learning
以对象为中心的视觉表示和强化学习
批准号:
2722103
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Abstract We will develop new object-centric sequence models for vision, with the intention to improve data-efficiencyand robustness to out-of-distribution environments for video prediction and vision-based reinforcement learning.Aims And Objectives The first stage of the proposed research aims to combine temporal predictive coding with object centric learning and apply the resulting model to video prediction. This stage will aim to answer several questions:What requirements should an object-centric video model satisfy? How can these requirements be reflected in the designof OCTPC? How should object properties, instantaneous attribute variables, stochastic temporal evolution of attributevariables, and inter-object relationships be represented? How should object instances be represented differently to objecttypes? How precisely should the temporal predictive coding mechanism be specified such that it is capable of learningsufficiently long-term dependencies? What are the commonalities and differences between existing approaches to object centric learning, and how do these relate to the performance of such algorithms? How does OCTPC perform on a variety ofvideo-prediction benchmarks? Which design choices and hyperparameters have the greatest effect on performance? Howdoes the performance, behaviour, and learning efficiency of OCTPC compare with models which aren't object-centric, orwhich don't use predictive coding?1The second stage of the proposed research aims to explore the possible benefits of OCTPC in reinforcement learning.There are two primary motivations for doing so. Firstly, reinforcement learning effectively depends on being able to predictthe future, because the ultimate objective is to choose a policy which maximises expected long-term future reward. It isplausible that a using a performant video prediction model as a component in model-based reinforcement learning wouldallow an agent to make more accurate predictions of how changes in policy would affect future experience, and thereforehow the policy should be changed in order to maximise future reward. Secondly, there is a neuroscientific principle whichstates that "the processing function of neocortical modules is qualitatively similar in all neocortical regions... there isnothing intrinsically motor about the motor cortex, nor sensory about the sensory cortex." [13]. Therefore, if it is foundthat object-centric inductive biases are useful for video prediction, then it may be the case that similar inductive biasesare useful in policy representations as well. It would be interesting to compare such policy representations to existingwork in hierarchical reinforcement learning [17], and explore whether such representations can improve sample-efficiencyin reinforcement learning, and robustness to out-of-distribution environments.Novelty Of The Research Methodology To our knowledge, combining temporal predictive coding with object-centriclearning has not previously been explored. It is plausible that exploring this combination will provide valuable contributionsand insights to machine learning, while also providing value to the cognitive sciences by moving closer to an understandingof human intelligence on the algorithmic level.Alignment To EPSRC's Strategies And Research Areas This research proposal aligns with the areas of Artificialintelligence and robotics theme, Artificial intelligence technologies, and Image and vision computing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金