Deep Spatial/temporal-level feature engineering for Tennis-based action recognition

Deep Spatial/temporal-level feature engineering for Tennis-based action recognition
复制标题

DOI:
10.1016/j.future.2021.06.022
复制
发表时间:
2021-07-01
影响因子:
7.5
通讯作者:
Na, Liu
Na, Liu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ning, Bai;Na, Liu

文献摘要

被引文献

相似文献

网球在全世界已成为一项越来越受欢迎的运动。近年来,基于3D视频的网球运动识别越来越受到人们的关注。该算法考虑了运动的时序信息,能够在时间层面上解决人体运动的不确定性。随着训练样本的增加,效率会相应降低。本文提出了一个基于动作标准序列的网球动作识别框架。通过特征提取将三维动作视频样本纳入动作序列,将动作标准序列编码为动态时间归一化度量下的序列平均优化问题。利用动态时间归一化重心平均算法(DBA)来解决这一问题。对于动作类别差异显著的网球场景,我们研究了多动作的标准序列学习,并据此提出了一种用于无监督学习的DBA-K-means聚类算法。在此基础上,提出了一种结合特征优化和图像相似度的人类网球动作识别方法。对比主成分分析(PCA)、PCA+ Pearson和PCA+ Spearman三种降维方法,发现PCA+ Pearson相关系数降维效果最好。同时,将全局特征八星模型与局部特征HOG特征进行降维后结合,充分表征人体运动。计算图像两两相邻帧之间的相似度。自适应地分配一个判别周期内单帧SVM分类结果的统计权值,最后对人体姿态识别结果进行两次分类。在标准数据集KTH上的实验表明,该算法的识别准确率为94.5%,优于其他方法。在视频人体动作识别领域具有很好的应用价值。实验结果表明,该方法可以进一步提高动作识别的效率和准确性。有效的特征提取有利于提高后续人体动作识别的准确性。(C) 2021 Elsevier B.V.版权所有
Tennis has becoming an increasingly popular sport throughout the world. Tennis motion recognition based on 3D video has attracted more and more attention in recent years. The algorithm based on dynamic time warping takes into account the timing sequence information of movements and can solve the uncertainty of human movement at temporal level. By increasing the training samples, the efficiency will decrease accordingly. This work presents a tennis action recognition framework based on action standard sequence. The 3D action video samples are incorporated into action sequences by feature extraction, wherein the action standard sequences are encoded as a sequence averaging optimization problem under the dynamic time normalization metric. The dynamic time normalization barycenter averaging algorithm (DBA) is leveraged to solve this problem. For the tennis scenery with significant differences in the action categories, we study the standard sequence learning of multiple actions, and accordingly propose a DBA-K-means clustering algorithm for unsupervised learning. Herein, a human tennis action recognition by integrating feature optimization and image similarity is proposed. The three dimensional reduction methods, including principal component analysis (PCA),PCA + Pearson, and PCA+ Spearman, were compared to prove that PCA+ Pearson correlation coefficient had the best dimensional reduction effect. Meanwhile, the global feature eight-star model is combined with the local feature HOG feature after dimensionally reduced to fully represent human movements. The similarity between pairwise adjacent frames of images was calculated. The statistical weight of single frame SVM classification results within a discriminant period is adaptively allocated, and finally the body pose recognition results are classified twice. Experiments on standard data set KTH show that the recognition accuracy of this algorithm is 94.5%, which is better than other methods. It has a good application value in the field of video human motion recognition. Also, we have demonstrated that this method can further improve the efficiency and accuracy of action recognition. Effective feature extraction is beneficial to improve the accuracy of subsequent human action recognition. (C) 2021 Elsevier B.V. All rights reserved.