Three-Stream Graph Convolutional Networks for Zero-Shot Action Recognition

Three-Stream Graph Convolutional Networks for Zero-Shot Action Recognition
复制标题

DOI:
10.1109/scisisis50064.2020.9322783
复制
发表时间:
2020-12
期刊:
2020 Joint 11th International Conference on Soft Computing and Intelligent Systems and 21st International Symposium on Advanced Intelligent Systems (SCIS-ISIS)
影响因子:
--
通讯作者:
Na Wu;K. Kawamoto
Na Wu;K. Kawamoto
中科院分区:
其他
文献类型:
--
作者:
Na Wu;K. Kawamoto

文献摘要

相似文献

最近在行动确认方面的发展导致行动类别的数目增加。为了适应这种增长,动作数据集需要大量昂贵且费力的注释视频。因此,零射击动作识别(ZSAR)变得越来越重要。目前,ZSAR方法主要有两种:一种是使用视频RGB图像数据,另一种是使用人体骨架数据。传统方法只使用其中一种类型的数据,而忽略其他数据,从而降低了模型的准确性。在本文中,我们提出了一个处理这两种类型数据的三流图卷积网络。我们对RGB数据使用两流图卷积网络,对骨架数据使用运动分支。将这两个输出与加权和相结合,我们的模型预测ZSAR的最终结果。通过在UCF101数据集上的实验,我们证明了我们的模型比基线模型提供了更好的精度。
Recent developments in action recognition have resulted in an increase in the number of action categories. To accommodate the increase, an action dataset requires a large number of expensive and laboriously annotated videos. Thus, zero-shot action recognition (ZSAR) has become increasingly important. At present, there are two main ZSAR methods: one uses the video RGB image data, and the other uses the skeleton data of the human body. Conventional approaches use only one of these types of data and ignore the other data, thereby reducing the model accuracy. In this paper, we propose a three-stream graph convolutional network that processes both types of data. We use a two-stream graph convolutional network for RGB data and a motion branch for skeleton data. Combining these two outputs with a weighted sum, our model predicts final results for ZSAR. With experiments on the dataset UCF101, we show that our model provides better accuracy than a baseline model.