3D Pose Tracking With Multitemplate Warping and SIFT Correspondences

3D Pose Tracking With Multitemplate Warping and SIFT Correspondences
复制标题

DOI:
10.1109/tcsvt.2015.2452782
复制
发表时间:
2016-11
影响因子:
8.4
通讯作者:
Shu Chen;Luming Liang;Wenzhang Liang;H. Foroosh
Shu Chen;Luming Liang;Wenzhang Liang;H. Foroosh
中科院分区:
工程技术1区
文献类型:
--
作者:
Shu Chen;Luming Liang;Wenzhang Liang;H. Foroosh

文献摘要

被引文献

相似文献

模板变形由于其适用于单目视频序列的灵活性,在基于视觉的3D运动跟踪和3D姿态估计中是一种流行的技术。然而,该方法受到两个主要限制,阻碍其在实践中的成功使用。首先,它需要在应用该方法之前校准相机。其次,如果框架间位移太大,则可能无法提供良好的结果。为了克服第一个问题,我们建议估计未知的焦距的相机从几个初始帧的迭代优化过程。为了缓解第二个问题,我们提出了一种跟踪方法的基础上结合密集光流和跟踪尺度不变特征变换(SIFT)功能提供的互补信息。虽然光流适用于小位移并提供准确的局部信息,但跟踪的SIFT特征在处理较大位移或全局变换方面更好。为了联合收割机这两条互补信息,我们引入遗忘因子来引导SIFT特征提供的3D姿态估计,并使用光流来细化最终结果。实验在三个公共数据库上进行,即,Biwi Head Pose数据集、BU数据集和麦吉尔Faces数据集。结果表明,所提出的解决方案提供了更准确的结果比基线方法,仅依赖于模板变形或SIFT功能。此外,该方法可以应用于各种各样的场景,由于规避了摄像机校准的需要,从而提供了一个更灵活的解决方案,比现有的方法的问题。
Template warping is a popular technique in vision-based 3D motion tracking and 3D pose estimation due to its flexibility of being applicable to monocular video sequences. However, the method suffers from two major limitations that hamper its successful use in practice. First, it requires the camera to be calibrated prior to applying the method. Second, it may fail to provide good results if the inter-frame displacements are too large. To overcome the first problem, we propose to estimate the unknown focal length of the camera from several initial frames by an iterative optimization process. To alleviate the second problem, we propose a tracking method based on combining complementary information provided by dense optical flow and tracked scale-invariant feature transform (SIFT) features. While optical flow is good for small displacements and provides accurate local information, tracked SIFT features are better at handling larger displacements or global transformations. To combine these two pieces of complementary information, we introduce a forgetting factor to bootstrap the 3D pose estimates provided by SIFT features, and refine the final results using optical flow. Experiments are performed on three public databases, i.e., the Biwi Head Pose dataset, the BU dataset, and the McGill Faces datasets. The results illustrate that the proposed solution provides more accurate results than baseline methods that rely solely on either template warping or SIFT features. In addition, the approach can be applied in a larger variety of scenarios, due to circumventing the need for camera calibration, thus providing a more flexible solution to the problem than existing methods.