Semantic structure from motion with object and point interactions

Semantic structure from motion with object and point interactions
复制标题

来自对象和点交互的运动的语义结构

DOI:
--
复制
发表时间:
2011
期刊:
IEEE International Conference on Computer Vision
影响因子:
--
通讯作者:
S. Savarese
S. Savarese
中科院分区:
--
文献类型:
--
作者:
Sid Ying;M. Bagra;S. Savarese

文献摘要

被引文献

相似文献

我们提出了一种新的方法,用于联合检测对象和恢复场景的几何形状(相机姿态,对象和场景点的3D位置)从多个半校准图像(相机内部参数是已知的)。为了实现这一任务,我们的方法模型高层次的语义(即对象类标签和相关特性,如位置和姿态)和对象和特征点在同一视图和跨视图的相互作用(相关性)。我们使用两个公共数据集-福特汽车数据集和Kinect Office数据集[1] -对最先进的基线方法验证了我们的算法,并表明我们:i)与基于点的SFM算法相比,显着提高了相机姿态估计结果; ii)与单独使用单个图像相比,实现了更好的2D和3D对象检测精度。我们的算法是至关重要的,在许多应用场景,包括对象操作和自主导航。
We propose a new method for jointly detecting objects and recovering the geometry of the scene (camera pose, object and scene point 3D locations) from multiple semi-calibrated images (camera internal parameters are known). To achieve this task, our method models high level semantics (i.e. object class labels and relevant characteristics such as location and pose) and the interaction (correlations) of objects and feature points within the same view and across views. We validate our algorithm against state-of-the-art baseline methods using two public datasets - Ford Car dataset and Kinect Office dataset [1] - and show that we: i) significantly improve the camera pose estimation results compared to point-based SFM algorithm; ii) achieve better 2D and 3D object detection accuracy than using single images separately. Our algorithm is critical in many application scenarios including object manipulation and autonomous navigation.