Human-Centric Scene Understanding from Single View 360 Video

Human-Centric Scene Understanding from Single View 360 Video
复制标题

DOI:
10.1109/3dv.2018.00046
复制
发表时间:
2018-09
期刊:
2018 International Conference on 3D Vision (3DV)
影响因子:
--
通讯作者:
Sam Fowler;Hansung Kim;A. Hilton
Sam Fowler;Hansung Kim;A. Hilton
中科院分区:
其他
文献类型:
--
作者:
Sam Fowler;Hansung Kim;A. Hilton

文献摘要

相似文献

在本文中,我们提出了一种方法,室内场景的理解,从观察的人在单视图球形视频。作为输入,我们的方法需要一个位于中心的球形视频捕获的室内场景,估计在整个长期捕获的3D本地化的人类行为。这项工作的核心贡献是在合成数据集上训练的深度卷积编码器-解码器网络,以从捕获的人类活动中重建启示区域。然后应用预测的示能表示分割来构成完整3D场景的重构,将示能表示分割集成到3D空间中。人类活动和示能分割之间的映射表明,对人类活动的全方位观察可以应用于3D重建等场景理解任务。我们表明,我们的方法只使用观察的人表现良好,对以前的方法,允许重建的闭塞区域和标签的场景启示。
In this paper, we propose an approach to indoor scene understanding from observation of people in single view spherical video. As input, our approach takes a centrally located spherical video capture of an indoor scene, estimating the 3D localisation of human actions performed throughout the long term capture. The central contribution of this work is a deep convolutional encoder-decoder network trained on a synthetic dataset to reconstruct regions of affordance from captured human activity. The predicted affordance segmentation is then applied to compose a reconstruction of the complete 3D scene, integrating the affordance segmentation into 3D space. The mapping learnt between human activity and affordance segmentation demonstrates that omnidirectional observation of human activity can be applied to scene understanding tasks such as 3D reconstruction. We show that our approach using only observation of people performs well against previous approaches, allowing reconstruction of occluded regions and labelling of scene affordances.