Rate-Accuracy Trade-Off in Video Classification with Deep Convolutional Neural Networks

Rate-Accuracy Trade-Off in Video Classification with Deep Convolutional Neural Networks
复制标题

DOI:
10.1109/icip.2018.8451666
复制
发表时间:
2018-10
期刊:
2018 25th IEEE International Conference on Image Processing (ICIP)
影响因子:
--
通讯作者:
Alhabib Abbas;Aaron Chadha;Y. Andreopoulos;M. Jubran
Alhabib Abbas;Aaron Chadha;Y. Andreopoulos;M. Jubran
中科院分区:
其他
文献类型:
--
作者:
Alhabib Abbas;Aaron Chadha;Y. Andreopoulos;M. Jubran

文献摘要

被引文献

相似文献

先进的视频分类系统对视频帧进行解码,以获得所需的纹理和运动表示,以供时空深度卷积神经网络(CNN)摄取和分析。然而,当考虑可视化物联网应用、监控系统和大型视频库的语义爬虫时,压缩视频内容和基于CNN的语义分析部分往往不会位于同一位置。这就需要通过网络传输压缩视频,并在带宽和能源消耗方面产生大量开销,从而极大地破坏了此类系统的部署潜力。在本文中,我们研究了基于CNN的视频分类在AVC/H.264编码视频的编码比特率和可达到的精度之间的权衡。与整个压缩视频比特流不同,我们只以显著降低的比特率保留运动向量和选定的纹理信息。基于两个CNN结构和两个动作识别数据集,在对分类精度影响不大的情况下,实现了38%-59%的码率节约。在两个CNN之间的简单的基于速率的选择表明,在准确度优雅地下降的情况下,甚至可以进一步节省比特率。这可以允许在网络上进行速率/精度优化的基于CNN的视频分类。
Advanced video classification systems decode video frames to derive the necessary texture and motion representations for ingestion and analysis by spatio-temporal deep convolutional neural networks (CNNs). However, when considering visual Internet -of- Things applications, surveillance systems and semantic crawlers of large video repositories, the compressed video content and the CNN-based semantic analysis parts do not tend to be co-located. This necessitates the transport of compressed video over networks and incurs significant overhead in bandwidth and energy consumption, thereby significantly undermining the deployment potential of such systems. In this paper, we investigate the trade-off between the encoding bitrate and the achievable accuracy of CNN-based video classification that ingests AVC/H.264 encoded videos. Instead of entire compressed video bitstreams, we only retain motion vector and selected texture information at significantly reduced bitrates. Based on two CNN architectures and two action recognition datasets, we achieve 38%-59% saving in bitrate with marginal impact in classification accuracy. A simple rate-based selection between the two CNNs shows that even further bitrate savings are possible with graceful degradation in accuracy. This may allow for rate/accuracy-optimized CNN-based video classification over networks.