SOSD-Net: Joint semantic object segmentation and depth estimation from monocular images

SOSD-Net: Joint semantic object segmentation and depth estimation from monocular images
复制标题

DOI:
10.1016/j.neucom.2021.01.126
复制
发表时间:
2021-03-05
期刊:
影响因子:
6
通讯作者:
Zhou, Jie
Zhou, Jie
中科院分区:
计算机科学2区
文献类型:
--
作者:
He, Lei;Lu, Jiwen;Zhou, Jie

文献摘要

被引文献

相似文献

深度估计和语义分割在场景理解中起着至关重要的作用。最先进的方法采用多任务学习,在像素级别同时学习这两个任务的模型。它们通常专注于共享公共特征或缝合来自相应分支的特征图。然而,这些方法缺乏深入考虑的相关性的几何线索和场景解析。本文首先通过对成像过程的分析,引入语义对象性的概念,利用这两个任务之间的几何关系,提出了一种基于对象性假设的语义对象分割和深度估计网络(SOSD-Net)。据我们所知,SOSD-Net是第一个利用几何约束同时进行单目深度估计和语义分割的网络。此外,考虑到这两个任务之间的相互隐含关系,我们利用期望最大化算法中的迭代思想来更有效地训练所提出的网络。在Cityscapes和NYU v2数据集上的大量实验结果证明了该方法的上级性能。(c)2021爱思唯尔有限公司版权所有。
Depth estimation and semantic segmentation play essential roles in scene understanding. The state-of -the-art methods employ multi-task learning to simultaneously learn models for these two tasks at the pixel-wise level. They usually focus on sharing the common features or stitching feature maps from the corresponding branches. However, these methods lack in-depth consideration on the correlation of the geometric cues and the scene parsing. In this paper, we first introduce the concept of semantic object-ness to exploit the geometric relationship of these two tasks through an analysis of the imaging process, then propose a Semantic Object Segmentation and Depth Estimation Network (SOSD-Net) based on the objectness assumption. To the best of our knowledge, SOSD-Net is the first network that exploits the geometry constraint for simultaneous monocular depth estimation and semantic segmentation. In addi-tion, considering the mutual implicit relationship between these two tasks, we exploit the iterative idea from the expectation-maximization algorithm to train the proposed network more effectively. Extensive experimental results on the Cityscapes and NYU v2 dataset are presented to demonstrate the superior performance of the proposed approach.(c) 2021 Elsevier B.V. All rights reserved.