Large-Scale Gesture Recognition With a Fusion of RGB-D Data Based on Saliency Theory and C3D Model

Large-Scale Gesture Recognition With a Fusion of RGB-D Data Based on Saliency Theory and C3D Model
复制标题

基于显着性理论和C3D模型的RGB-D数据融合的大规模手势识别

DOI:
10.1109/tcsvt.2017.2749509
复制
发表时间:
2018-10-01
影响因子:
8.4
通讯作者:
Song, Jianfeng
Song, Jianfeng
中科院分区:
工程技术1区
文献类型:
--
作者:
Li, Yunan;Miao, Qiguang;Song, Jianfeng

文献摘要

被引文献

相似文献

手势识别由于其广泛的应用在计算机视觉领域引起了广泛的关注。然而,基于视频的大规模手势识别任务仍然面临着许多挑战,因为许多手势无关的因素,如背景可能会干扰识别的准确性。为了更好地识别大规模视频中的手势,本文提出了一种基于RGB-D数据的方法,其中“RGB-D”是指由Kinect等特定设备同时捕获的RGB和深度数据。为了更好地学习手势细节,我们首先使用自适应帧统一策略来统一输入的帧数,然后将RGB和深度数据发送到C3 D模型以分别提取时空特征。为了减少手势无关因素的干扰,显著性理论也被用来产生辅助数据。然后将这些数据的特征进行组合以提高性能,这也可以避免不合理的合成数据,因为C3 D特征的维度是均匀的。最后对几种分类器的性能进行了测试,并从中选出最佳的SVM分类器输出最终的分类精度。我们的方法在Chalearn的IsoGD验证和测试子集上分别实现了52.04%和59.43%的准确率,两者都优于我们在ICPR 2016中报告的Chalearn大规模手势识别挑战赛中的结果。
Gesture recognition has raised wide attention in computer vision owing to its many applications. However, the task of video-based large-scale gesture recognition yet faces many challenges, since many gesture-irrelevant factors like the background may disturb the recognition accuracy. To better recognize gestures with large-scale videos, we propose a method based on RGB-D data in this paper, where the “RGB-D” means RGB and depth data captured simultaneously by specific devices like Kinect. To learn gesture details better, we first use an adaptive frame unification strategy to unify the frame number of inputs, and then the RGB and depth data are sent to the C3D model to extract spatiotemporal features, respectively. In order to alleviate the interference of gesture-irrelevant factors, the saliency theory is also employed to generate auxiliary data. Next the features of these data are combined to boost the performance, which can also avoid unreasonable synthetic data, since the dimension of C3D features is uniform. Finally the performances of several classifiers are tested and the best one of SVM classifier is selected to output the ultimate accuracy. Our approach achieves 52.04% and 59.43% accuracy on the validation and testing subset of the Chalearn LAP IsoGD, respectively, both of which outperform our results in the chalearn LAP Large-scale Gesture Recognition Challenge as reported in ICPR 2016.