课题基金 / 基金详情

Self-supervised Monocular Depth Estimation

Self-supervised Monocular Depth Estimation
自监督单目深度估计
批准号:
2747408
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
1.1单目深度估计计算机视觉的一项基本任务是了解我们和其他物体在空间中的3D定位。许多先前的深度估计方法使用立体深度,虽然这是一种有效的技术,但最近深度神经网络的进步导致了单目深度估计的发展。单目深度估计从单个RGB图像产生幻觉深度。传统上,监督方法将被用于确定每像素图像的深度作为回归任务。然而,最近的发展使用了利用连续RGB帧之间的视图合成的自我监督方法。这些方法利用自位估计和深度估计将图像帧反向扭曲到其连续的图像帧。这导致的能力,不仅了解一个场景的深度,但了解相机的运动通过这个场景。这些方法已经显示出有希望的结果,并在不断改进,但仍有许多进步要做。1.2改进现在是一个众所周知的问题,这些方法与动态对象斗争,因为它们在所有场景中都假定为刚性运动。大多数最新的方法试图通过允许系统忽略导致损失函数中误差的区域来避免这些下降。在这个项目中,我们的目标是利用这些不便来教导深度和自我姿态网络更多地了解场景。此外,通过关注这些动态物体,我们可以以自我监督的方式跟踪车辆和行人的运动,这是自动驾驶的重要任务。此外,有大量的关注场景,只考虑日光视频与大部分晴朗的天气。以今天的SoTA模型为例,我们看到它们在雨、雪、雾和其他更极端的天气事件中挣扎,每当天气不晴朗时,就会导致故障。考虑到2021年英国有149天的降雨,我们的深度和姿势模型处理降雨显然是一个问题。现代方法试图通过微调不同天气事件的模型来处理这个问题,导致每种天气条件下的单独模型,这表明当前的深度模型不能很好地泛化。我们的目标是改进这种方法,同时使网络具有更大的通用性。由于这些方法是病态的,因此对深度和姿态网络进行某种形式的不确定性估计是有益的。目前在该领域的工作是有限的,不确定性估计工作需要发展为自我姿态和其他目标姿态估计。这一点以及对深度不确定性的进一步改进将有直接的工业应用,这将是本项目的重点。最后,这些方法难以处理无纹理的区域,这在以前的工作中已经得到了解决。为了进一步完善这些方法,我们将在无纹理区域中使用更有效的损失函数进行视图合成。最后,我们的目标是处理所有这些问题,同时减少模型的计算使用量,使这些方法能够应用于实际应用。1.3行业应用大多数自动驾驶汽车使用雷达、Li-DAR和立体传感器的传感器融合设置。这种设置产生精确的3D重建,但导致显著的物理成本。这种传感器组合可能会导致不准确的结果,因为传感器可能不一致。最终,在车辆周围安装一个简单、精确的单目摄像头来取代这种配置,将减少车队中每辆车对Li-DAR和立体摄像头的需求。虽然这种改进对于单个车辆来说是适度的,但对于大型车队来说,这将显著降低成本。”
英文摘要
"1.1 Monocular Depth EstimationA fundamental task in computer vision is understanding our and other objects'3D positioning in space. Many prior methods of depth estimation made use of stereoscopic depth, while this was an effective technique, recent advancements in Deep Neural Networks have led to the development of monocular depth estimation. Monocular depth estimation hallucinates depth from a single RGB image. Traditionally, supervised methods would have been used to determine the depth of an image per pixel as a regression task. However, recent developments use self-supervising methods that take advantage of view-synthesis between consecutive RGB frames. These methods make use of ego-pose estimation, and depth estimation to inverse warp an image frame to its consecutive image frame. This leads to the ability to not only understand the depth of a scene but to understand the camera'smotion through this scene. These methods have shown promising results and are constantly improving but there are many advancements to be made.1.2 ImprovementsIt is now a well-known issue that these methods struggle with dynamic objects, as they assume rigid motion in all scenes. Most recent methods attempt to avoid these downfalls by allowing the system to ignore the regions that cause the error in the loss function. In this project, we aim to employ these inconveniences to teach the depth and ego-pose networks more about the scene. Also, by focusing on these dynamic objects, we can track the motionof vehicles and pedestrians in a self-supervised manner, which is a vital task for autonomous driving. Furthermore, there has been a large focus on scenes that only take into consideration daylight videos with mostly sunny clear weather. Taking today's SoTA models we see that they struggle with rain, snow, fog and other more extreme weather events leading to failure cases whenever the weather is not clear. Given that in 2021 we had 149 days of rain in the UK it is a clear issue for our depth and pose models to handle rain. Modern methods attempt to handle this by fine-tuning the models for different weather events, leading to separate models for each weather condition, which demonstrates that the current depth models do not generalise well. We aim to improve this methodology while leading to much greater generalisability of the networks. As these methods are ill-posed, it is beneficial for these methods to have some form of uncertainty estimations for the depth and pose networks. Current work in this area is limited, and uncertainty estimation work needs to be developed for ego-pose and other object-pose estimations. This and further improvements for depth uncertainty would have direct industrial applications and would be a focus of this project.Penultimately, these methods struggle with texture-less regions, which has been addressed in previous work. To further these methods, we will use a more efficient loss function for view synthesis in textureless regions. Finally, we aim to handle all of these issues while leading to reductions in computation usage of the models, allowing for these methods to be applied to practical applications.1.3 Industry ApplicationMost self-driving vehicles use a sensor fusion setup of Radar, Li-DAR and stereo sensors. This setup generates accurate 3D reconstructions but leads to significant physical costs. This combination of sensors can suffer in inaccurate results as sensors may disagree. Ultimately, replacing this configuration with a simple, accurate, monocular camera setup around the vehicle would reduce the need for Li-DAR and stereo cameras for each vehicle in a fleet. While this improvement is moderate for a single vehicle, on a large fleet this would lead to compelling reductions in costs."
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于指点触控行为的身份认证与监控方法研究
  • 批准号:
    61175039
  • 项目类别:
    面上项目
  • 资助金额:
    59.0万元
  • 批准年份:
    2011
  • 负责人:
    蔡忠闽
  • 依托单位: