Unsupervised cataloguing of objects in a 3D scene
Unsupervised cataloguing of objects in a 3D scene
批准号:
2577387
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
对研究背景的简要描述,包括潜在的影响:很难或有时不可能获得足够大量的标记或甚至原始数据来训练3D感知算法。另一方面,视频作为丰富的信息来源,提供了毫不费力地捕捉场景和物体的全部3D复杂性的潜力,因此成为了应对这一挑战的冷门解决方案。在这种背景下,利用视频数据的方便性和无处不在的特点,自动提取和索引3D场景中的对象是本研究的重点。这个项目的目标是建立算法,能够推理高级概念,如“客观性”,以发现场景中的对象,同时也能够重建这些对象并预测它们在场景的3D模型中的精确布局。如果成功,我们的系统可以用来从单目视频生成3D场景的虚拟模型,这些模型可以用来创建逼真的增强现实体验。此外,该系统还具有直接的现实应用--例如,跟踪零售系统中的库存或帮助仓库中的库存管理。主要目标是开发一种能够自动识别3D场景中的对象并对其进行唯一索引的系统。除了唯一索引对象外,我们的系统还应该能够准确地重建对象并预测它们在场景的3D模型中的精确布局。我们的目标是以一种无人监督的方式训练这个系统,而不是明确使用3D数据或注释,因为它们很难获得。取而代之的是,我们将从更便宜的来源提取信息,例如捕捉场景的视频或多个视点的场景图像。该项目还将探索使用文本作为另一种形式,以提供对场景或对象的语义属性的弱监督。另一种可能性是在动态场景中工作,其中运动可以是识别场景中对象的有用线索。研究方法的新奇之处:1.识别和编目场景中的对象是一个具有挑战性的问题,因为我们没有预定义的对象列表,因此我们将开发算法来训练一个系统,该系统可以推理场景中的对象来发现对象。另一个困难来自编目方面,因为我们还需要验证新发现的对象是否已经与场景中已编目的对象相对应(S)。因此,我们的研究还将集中在能够有效地用于编目的3D形状感知实例检索算法上。由于很难获得3D注解,我们将探索如何结合其他监督来源,例如来自视频的文本或运动提示来训练我们的模型。与EPSRC的战略和研究领域(项目涉及的研究领域)图像和视觉计算(EPSRC链接)有关的任何公司或合作者:此DPhil项目的顾问将是牛津大学视觉几何组的Andrea Vedaldi、Andrew Zisserman、João Henrique和Iro Laina
英文摘要
Brief description of the context of the research including potential impact:It is difficult or sometimes impossible to obtain a sufficiently large amounts of labelled or evenraw data to train 3D perception algorithms. Videos, on the other hand, have emerged as apragmatic solution to this challenge as they serve as a rich source of information, offeringthe potential to effortlessly capture scenes and objects in their full 3D complexity. In this context,the focus of this research is the pivotal task of automated extraction and indexing of objectswithin 3D scenes, leveraging the convenience and ubiquity of video data. This project aims atbuilding algorithms that can reason about high-level concepts such as "objectness" to discoverobjects in a scene, while also being able to reconstruct these objects and predict their preciselayout in the 3D model of the scene. If successful, our system can be used to generate virtualmodels of 3D scenes from monocular videos, which could be used to create realistic augmentedreality experiences. Furthermore, this system also has direct real-world applications -- e.g., tokeep track of inventory in retail systems or to help with stock management in warehouses.Aims and Objectives1. The primary objective is to develop a system that can automatically identify objects in a3D scene and uniquely index them.2. In addition to uniquely indexing the objects, our system should also be able to accuratelyreconstruct the objects and predict their precise layout in the 3D model of the scene.3. We aim to train this system in an unsupervised manner and without explicitly using 3Ddata or annotations as they are expensive to obtain. We will instead extract informationfrom cheaper sources such as videos capturing the scene or images of the scene frommultiple viewpoints.4. The project will also explore using text as another modality to provide weak supervisionabout the semantic properties of the scene or objects. Another possibility is to work ondynamic scenes where motion can be a useful cue for identifying objects in the scene.Novelty of the research methodology:1. Identifying and cataloguing objects in a scene is a challenging problem because we donot possess a predefined list of objects, hence we will develop algorithms to train asystem that can reason about "objectness" to discover objects in the scene.2. Another difficulty comes from the cataloguing aspect as we also need to verify if a newlyfound object already corresponds to a catalogued object in the scene(s). Hence, ourresearch will also focus on 3D shape-aware instance retrieval algorithms that can beefficiently used for cataloguing.3. Since 3D annotations are difficult to obtain, we will explore ways to incorporate othersources of supervision such as text or motion cues from videos to train our models.Alignment to EPSRC's strategies and research areas (which EPSRC research area theproject relates to)Image and vision computing (EPSRC link)Any companies or collaborators involved: The advisors for this DPhil project will be AndreaVedaldi, Andrew Zisserman, João Henriques and Iro Laina from the Visual Geometry Group atUniversity of Oxford
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金