PWP3D: Real-Time Segmentation and Tracking of 3D Objects

PWP3D: Real-Time Segmentation and Tracking of 3D Objects
复制标题

DOI:
10.1007/s11263-011-0514-3
复制
发表时间:
2012-07
影响因子:
19.5
通讯作者:
V. Prisacariu;I. Reid
V. Prisacariu;I. Reid
中科院分区:
计算机科学2区
文献类型:
--
作者:
V. Prisacariu;I. Reid

文献摘要

被引文献

相似文献

我们制定了一个概率框架,同时基于区域的2D分割和2D到3D姿态跟踪,使用已知的3D模型。给定这样的模型,我们的目标是通过直接优化3D姿态参数来最大限度地提高统计前景和背景外观模型之间的区分度。前景区域由符号距离嵌入函数的零水平集划定,并且我们基于像素后验隶属度概率(而不是似然度)定义该区域及其紧邻背景环境的能量。我们推导出该能量相对于3D对象的姿态参数的微分,这意味着我们可以使用标准的基于梯度的非线性最小化技术来搜索正确的姿态。我们提出了新的增强在像素级的时间一致性和改进的在线外观模型自适应的基础上。此外,我们的方法的直接扩展导致多相机和多对象跟踪作为同一框架的一部分。在我们的算法中的大部分处理的并行性质意味着它是服从GPU加速,我们给出了我们的实时实现的细节,我们用它来生成实验结果的真实的和人工视频序列,与一些3D模型。这些实验证明了使用像素后验而不是似然的好处,并展示了我们的跟踪器的质量,例如对遮挡和运动模糊(以及一些故障模式)的鲁棒性。
We formulate a probabilistic framework for simultaneous region-based 2D segmentation and 2D to 3D pose tracking, using a known 3D model. Given such a model, we aim to maximise the discrimination between statistical foreground and background appearance models, via direct optimisation of the 3D pose parameters. The foreground region is delineated by the zero-level-set of a signed distance embedding function, and we define an energy over this region and its immediate background surroundings based on pixel-wise posterior membership probabilities (as opposed to likelihoods). We derive the differentials of this energy with respect to the pose parameters of the 3D object, meaning we can conduct a search for the correct pose using standard gradient-based non-linear minimisation techniques. We propose novel enhancements at the pixel level based on temporal consistency and improved online appearance model adaptation. Furthermore, straightforward extensions of our method lead to multi-camera and multi-object tracking as part of the same framework. The parallel nature of much of the processing in our algorithm means it is amenable to GPU acceleration, and we give details of our real-time implementation, which we use to generate experimental results on both real and artificial video sequences, with a number of 3D models. These experiments demonstrate the benefit of using pixel-wise posteriors rather than likelihoods, and showcase the qualities, such as robustness to occlusions and motion blur (and also some failure modes), of our tracker.