课题基金 / 基金详情

ROSSINI: Reconstructing 3D structure from single images: a perceptual reconstruction approach

ROSSINI: Reconstructing 3D structure from single images: a perceptual reconstruction approach
ROSSINI:从单个图像重建 3D 结构:感知重建方法
批准号:
EP/S016260/1
负责人:
Andrew Schofield
金额:
$52.23万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

Andrew Schofield的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Consumers enjoy the immersive experience of 3D content in cinema, TV and virtual reality (VR), but it is expensive to produce. Filming a 3D movie requires two cameras to simulate the two eyes of the viewer. A common but expensive alternative is to film a single view, then use video artists to create the left and right eyes' views in post-production. What if a computer could automatically produce a 3D model (and binocular images) from 2D content: 'lifting images into 3D'? This is the overarching aim of this project. Lifting into 3D has multiple uses, such as route planning for robots, obstacle avoidance for autonomous vehicles, alongside applications in VR and cinema.Estimating 3D structure from a 2D image is difficult because in principle, the image could have been created from an infinite number of 3D scenes. Identifying which of these possible worlds is correct is very hard, yet humans interpret 2D images as 3D scenes all the time. We do this every time we look at a photograph, watch TV or gaze into the distance, where binocular depth cues are weak. Although we make some errors in judging distances, our ability to quickly understand the layout of any scene enables us to navigate through and interact with any environment.Computer scientists have built machine vision systems for lifting to 3D by incorporating scene constraints. A popular technique is to train a deep neural network with a collection of 2D images and associated 3D range data. However, to be successful, this approach requires a very large dataset, which can be expensive to acquire. Furthermore, performance is only as good as the dataset is complete: if the system encounters a type of scene or geometry that does not conform to the training dataset, it will fail. Most methods have been trained for specific situations - e.g. indoor, or street scenes - and these systems are typically less effective for rural scenes and less flexible and robust than humans. Finally, such systems provide a single reconstructed output, without any measure of uncertainty. The user must assume that the 3D reconstruction is correct, which will be a costly assumption in many cases.Computer systems are designed and evaluated based upon their accuracy with respect to the real world. However, the ultimate goal of lifting into 3D is not perfect accuracy - rather it is to deliver a 3D representation that provides a useful and compelling visual experience for a human observer, or to guide a robot whilst avoiding obstacles. Importantly, humans are expert at interacting with 3D environments, even though our perception can deviate substantially from true metric depth. This suggests that human-like representations are both achievable and sufficient, in any and all environments.ROSSINI will develop a new machine vision system for 3D reconstruction that is more flexible and robust than previous methods. Focussing on static images, we will identify key structural features that are important to humans. We will combine neural networks with computer vision methods to form human-like descriptions of scenes and 3D scene models. Our aims are to (i) produce 3D representations that look correct to humans even if they are not strictly geometrically correct (ii) do so for all types of scene and (iii) express the uncertainty inherent in each reconstruction. To this end we will collect data on human interpretation of images and incorporate this information into our network. Our novel training method will learn from humans and existing ground truth datasets; the training algorithm selecting the most useful human tasks (i.e. judge depth within a particular image) to maximise learning. Importantly, the inclusion of human perceptual data should reduce the overall quantity of training data required, while mitigating the risk of over-reliance on a specific dataset. Moreover, when fully trained, our system will produce 3D reconstructions alongside information about the reliability of the depth estimates.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/wacvw58289.2023.00069
发表时间: 2022-11
期刊: 2023 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW)
影响因子: --
作者: [Jaime Spencer;C. Qian;Chris Russell;Simon Hadfield;E. Graf;W. Adams;A. Schofield;J. Elder;R. Bowden;Heng Cong;S. Mattoccia;Matteo Poggi;Zeeshan Khan Suri;Yang Tang;Fabio Tosi;Hao Wang;Youming Zhang;Yusheng Zhang;Chaoqiang Zhao]
通讯作者: Jaime Spencer;C. Qian;Chris Russell;Simon Hadfield;E. Graf;W. Adams;A. Schofield;J. Elder;R. Bowden;Heng Cong;S. Mattoccia;Matteo Poggi;Zeeshan Khan Suri;Yang Tang;Fabio Tosi;Hao Wang;Youming Zhang;Yusheng Zhang;Chaoqiang Zhao
Surface Attitude Judgements in monocular and stereo textures: a method evaluation
单目和立体纹理的表面姿态判断:方法评估
DOI: --
发表时间: 2021
期刊:
影响因子: --
作者: [Qian CS]
通讯作者: Qian CS
DOI: 10.1016/j.visres.2023.108275
发表时间: 2023
期刊: Vision research
影响因子: 1.8
作者: [Skog E]
通讯作者: Skog E
DOI: 10.1109/cvprw59228.2023.00308
发表时间: 2023-04
期刊: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子: --
作者: [Jaime Spencer;C. Qian;Michaela Trescakova;Chris Russell;Simon Hadfield;E. Graf;W. Adams;A. Schofield;J. Elder;R. Bowden;Ali Anwar;Hao Chen;Xiaozhi Chen;Kai Cheng;Yuchao Dai;Huynh Thai Hoa;Sadat Hossain;Jian-qiang Huang;Mohan Jing;Bo Li;Chao Li;Baojun Li;Zhiwen Liu;S. Mattoccia;Siegfried Mercelis;Myungwoo Nam;Matteo Poggi;Xiaohua Qi;Jiahui Ren;Yang Tang;Fabio Tosi;L. Trinh;S M Nadim Uddin;Khan Muhammad Umair;Kaixuan Wang;Yufei Wang;Yixing Wang;Mochu Xiang;Guangkai Xu;Wei Yin;Jun Yu;Qi Zhang;Chaoqiang Zhao]
通讯作者: Jaime Spencer;C. Qian;Michaela Trescakova;Chris Russell;Simon Hadfield;E. Graf;W. Adams;A. Schofield;J. Elder;R. Bowden;Ali Anwar;Hao Chen;Xiaozhi Chen;Kai Cheng;Yuchao Dai;Huynh Thai Hoa;Sadat Hossain;Jian-qiang Huang;Mohan Jing;Bo Li;Chao Li;Baojun Li;Zhiwen Liu;S. Mattoccia;Siegfried Mercelis;Myungwoo Nam;Matteo Poggi;Xiaohua Qi;Jiahui Ren;Yang Tang;Fabio Tosi;L. Trinh;S M Nadim Uddin;Khan Muhammad Umair;Kaixuan Wang;Yufei Wang;Yixing Wang;Mochu Xiang;Guangkai Xu;Wei Yin;Jun Yu;Qi Zhang;Chaoqiang Zhao
8
    Visual Image Interpretation in Humans and Machines
    • 批准号:
      EP/L014564/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $15.46万
    • 财政年份:
      2014
    • 负责人:
      Andrew Schofield
    • 依托单位:
    Beyond Luttinger Liquids-spin-charge separation at high excitation energies
    • 批准号:
      EP/J016888/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $27.85万
    • 财政年份:
      2012
    • 负责人:
      Andrew Schofield
    • 依托单位:
    Estimating the intrinsic characteristics of real images to aid analysis
    • 批准号:
      EP/F026269/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $48.57万
    • 财政年份:
      2008
    • 负责人:
      Andrew Schofield
    • 依托单位:
    Verification of Soil Liquefaction Analysis by Coordinated Geotechnical Centrifuge Studies
    • 批准号:
      9000927
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $15.98万
    • 财政年份:
      1989
    • 负责人:
      Andrew Schofield
    • 依托单位:
    海外基金