Scene Descriptor Expressing Ambiguity in Information Recovery Based on Incomplete Partial Observation

Scene Descriptor Expressing Ambiguity in Information Recovery Based on Incomplete Partial Observation
复制标题

DOI:
10.1109/iros51168.2021.9636576
复制
发表时间:
2021-09
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Takaaki Fukui;T. Matsuo;N. Shimada
Takaaki Fukui;T. Matsuo;N. Shimada
中科院分区:
其他
文献类型:
--
作者:
Takaaki Fukui;T. Matsuo;N. Shimada

文献摘要

相似文献

在最近的研究中,深度学习的广泛应用使得各种大规模的图像数据集成为可能,它使得基于图像的三维场景重建的性能得到了提高。一些研究在摄像机参数和图像地标对应关系未知的情况下,通过将来自训练数据集的先验知识与所获得的部分观测相结合,来估计包括遮挡或不可见部分在内的整个3D场景与所获得的部分观测相一致。提出了一种基于深度学习的场景重构框架,该框架能够直接表示和处理场景重构中的歧义。我们介绍了一种神经网络,它将目标场景编码为描述符。该网络将部分观测作为输入,并输出包含与给定观测一致的所有场景的场景描述符的参数集。由于没有几何信息或地标对应可用(定义不明确的情况),输入观测在某种意义上可能是不完整的,即它们没有足够的信息片段来唯一地确定整个场景。该网络基于完整的3-D场景和可能的局部观测的数据集进行训练,以便它可以从不完整的观测中预测未见的部分。本文介绍了一种在编解码器结构中引入描述符空间的方法,该方法采用了新的损失函数的定义来衡量“有效性”、“一致性”和“可重复性”。当获得同一三维场景的一系列部分和不完全观测时,通过将每个观测的描述符集参数积分到一个描述符集来显式地处理重建歧义。
In recent studies, the widespread of deep learning has made many kinds of large-scale image datasets available and it has enabled to improve the performance of image-based 3-D scene reconstruction. Several studies estimate whole 3-D scenes including occluded or unseen parts consistent with the obtained partial observations by integrating prior knowledge from training datasets with them, under no camera parameters nor image landmark correspondence are known. Although they generate “discrete” scene instances, they cannot represent and treat their “ambiguity” at all.This paper proposes a novel deep-learning-based framework that can directly represent and treat the ambiguity of scene reconstructions. We introduce a neural network which encodes a target scene as a descriptor. The network takes partial observations as input and outputs a parametric set of the scene descriptors containing all scenes consistent with given observations. The input observations may be “incomplete” in the sense that they do not have enough pieces of information to uniquely determine the whole scene due to neither geometry in-formation nor landmark correspondences available (ill-defined cases). The network is trained based on the dataset of the complete 3-D scenes and possible partial observations so that it can predict the unseen parts from incomplete observations. The paper introduces the method to induce such a descriptor space into the encoder/decoder architecture by employing novel definitions of loss functions measuring “validity”, “consistency” and “reproducibility”. When the series of partial and incom-plete observations for the same 3-D scene is obtained, the reconstruction ambiguity is explicitly treated by parametrically integrating the descriptor set for each observation into one descriptor set.