Three-dimensional model-based object recognition and segmentation in cluttered scenes

Three-dimensional model-based object recognition and segmentation in cluttered scenes
复制标题

DOI:
10.1109/tpami.2006.213
复制
发表时间:
2006-10-01
影响因子:
23.6
通讯作者:
Owens, Robyn
Owens, Robyn
中科院分区:
计算机科学1区
文献类型:
--
作者:
Mian, Ajmal S.;Bennamoun, Mohammed;Owens, Robyn

文献摘要

被引文献

相似文献

视点无关的自由形式的对象的识别和它们的分割中存在的杂波和遮挡是一个具有挑战性的任务。我们提出了一种新的3D模型为基础的算法,自动和有效地执行这项任务。一个物体的3D模型是从它的多个无序的范围图像(视图)中自动离线构建的。这些视图被转换为多维表表示(我们称之为张量)。通过使用基于哈希表的投票方案同时匹配视图的张量与其余视图的张量,在这些视图之间自动建立对应关系。这将产生一个相对变换图,用于在视图集成到无缝3D模型之前注册视图。这些模型及其张量表示构成了模型库。在在线识别过程中,来自场景的张量通过投票与库中的张量同时匹配。计算获得最多投票的模型张量的相似性度量。具有最高相似性的模型被转换到场景中,如果它与场景中的对象精确对齐,则该对象被声明为已识别并被分割。重复该过程,直到场景被完全分割。对由55个模型和610个场景组成的真实的和合成数据进行了实验,总体识别率达到95%。与自旋图像的比较表明,我们的算法在识别率和效率方面是上级的。
Viewpoint independent recognition of free-form objects and their segmentation in the presence of clutter and occlusions is a challenging task. We present a novel 3D model-based algorithm which performs this task automatically and efficiently. A 3D model of an object is automatically constructed offline from its multiple unordered range images (views). These views are converted into multidimensional table representations (which we refer to as tensors). Correspondences are automatically established between these views by simultaneously matching the tensors of a view with those of the remaining views using a hash table-based voting scheme. This results in a graph of relative transformations used to register the views before they are integrated into a seamless 3D model. These models and their tensor representations constitute the model library. During online recognition, a tensor from the scene is simultaneously matched with those in the library by casting votes. Similarity measures are calculated for the model tensors which receive the most votes. The model with the highest similarity is transformed to the scene and, if it aligns accurately with an object in the scene, that object is declared as recognized and is segmented. This process is repeated until the scene is completely segmented. Experiments were performed on real and synthetic data comprised of 55 models and 610 scenes and an overall recognition rate of 95 percent was achieved. Comparison with the spin images revealed that our algorithm is superior in terms of recognition rate and efficiency.