Evaluation of Triple-Stream Convolutional Networks for Action Recognition

Evaluation of Triple-Stream Convolutional Networks for Action Recognition
复制标题

DOI:
10.1109/dicta.2017.8227428
复制
发表时间:
2017-11
期刊:
2017 International Conference on Digital Image Computing: Techniques and Applications (DICTA)
影响因子:
--
通讯作者:
Dichao Liu;Yu Wang;Jien Kato
Dichao Liu;Yu Wang;Jien Kato
中科院分区:
其他
文献类型:
--
作者:
Dichao Liu;Yu Wang;Jien Kato

文献摘要

相似文献

近年来,双流卷积网络取得了显著的性能。特别是,通过捕捉外观和运动信息,时空双流网络带来了显著的改进。另一方面,动态图像作为视频的一种强有力的表现形式,也被证实可以为空间外观提供补充信息。受这些工作的启发,我们通过融合输入为动态图像的第三个网络流,提出了三流卷积网络。在本文中,我们实现了所提出的三流卷积网络,并从两个方面对它们进行了评估:(A)如何通过添加动态流来提高端到端的整体分类性能;(B)使用训练好的三流卷积网络来进行分类是有效的。我们的评估表明,与单一网络(空间和时间)和融合时空双流网络相比,该算法都有改进。
Recently, Two-Stream Convolutional Network has achieved remarkable performance. Especially, by capturing appearance and motion information, spatial-temporal two- stream networks bring noticeable improvement. On the other hand, dynamic image, which is a powerful representation for videos, has also been confirmed to provide complimentary information to spatial appearance. Inspired by these works, we proposed Triple-Stream Convolutional Networks by fusing a third network stream whose input is dynamic image. In this paper, we implement the proposed Triple-Stream Convolutional Networks and evaluated them in two aspects: (a) how the overall end-to-end classification performance can be benefited by adding the dynamic stream; (b) which way is efficient to use the trained Triple-Stream Convolutional Networks in classification. Our evaluation shows improvements over both single networks (spatial and temporal) and Fused Spatial-temporal Two-Stream Network.