Unsupervised cataloguing of objects in a 3D scene
Unsupervised cataloguing of objects in a 3D scene
批准号:
2577387
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Brief description of the context of the research including potential impact:It is difficult or sometimes impossible to obtain a sufficiently large amounts of labelled or evenraw data to train 3D perception algorithms. Videos, on the other hand, have emerged as apragmatic solution to this challenge as they serve as a rich source of information, offeringthe potential to effortlessly capture scenes and objects in their full 3D complexity. In this context,the focus of this research is the pivotal task of automated extraction and indexing of objectswithin 3D scenes, leveraging the convenience and ubiquity of video data. This project aims atbuilding algorithms that can reason about high-level concepts such as "objectness" to discoverobjects in a scene, while also being able to reconstruct these objects and predict their preciselayout in the 3D model of the scene. If successful, our system can be used to generate virtualmodels of 3D scenes from monocular videos, which could be used to create realistic augmentedreality experiences. Furthermore, this system also has direct real-world applications -- e.g., tokeep track of inventory in retail systems or to help with stock management in warehouses.Aims and Objectives1. The primary objective is to develop a system that can automatically identify objects in a3D scene and uniquely index them.2. In addition to uniquely indexing the objects, our system should also be able to accuratelyreconstruct the objects and predict their precise layout in the 3D model of the scene.3. We aim to train this system in an unsupervised manner and without explicitly using 3Ddata or annotations as they are expensive to obtain. We will instead extract informationfrom cheaper sources such as videos capturing the scene or images of the scene frommultiple viewpoints.4. The project will also explore using text as another modality to provide weak supervisionabout the semantic properties of the scene or objects. Another possibility is to work ondynamic scenes where motion can be a useful cue for identifying objects in the scene.Novelty of the research methodology:1. Identifying and cataloguing objects in a scene is a challenging problem because we donot possess a predefined list of objects, hence we will develop algorithms to train asystem that can reason about "objectness" to discover objects in the scene.2. Another difficulty comes from the cataloguing aspect as we also need to verify if a newlyfound object already corresponds to a catalogued object in the scene(s). Hence, ourresearch will also focus on 3D shape-aware instance retrieval algorithms that can beefficiently used for cataloguing.3. Since 3D annotations are difficult to obtain, we will explore ways to incorporate othersources of supervision such as text or motion cues from videos to train our models.Alignment to EPSRC's strategies and research areas (which EPSRC research area theproject relates to)Image and vision computing (EPSRC link)Any companies or collaborators involved: The advisors for this DPhil project will be AndreaVedaldi, Andrew Zisserman, João Henriques and Iro Laina from the Visual Geometry Group atUniversity of Oxford
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金