A long video caption generation algorithm for big video data retrieval

A long video caption generation algorithm for big video data retrieval
复制标题

DOI:
10.1016/j.future.2018.10.054
复制
发表时间:
2019-04-01
影响因子:
7.5
通讯作者:
Wan, Shaohua
Wan, Shaohua
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ding, Songtao;Qu, Shiru;Wan, Shaohua

文献摘要

被引文献

相似文献

人们拍摄的视频通常与他们生活中的某些重要时刻相关。但随着大数据时代的到来,检索和观看所需的时间可能会令人望而生畏。本文针对长视频分割的应用提出了新颖的技术,可以有效缩短检索时间。通过改进的时空兴趣点(STIPs)检测算法来检测长视频的运动程度。然后对过滤后的长视频进行超帧分割,得到感兴趣的长视频片段。在关键帧的选择中,利用视频片段上已经获得的STIP来构建感兴趣区域,并利用这些感兴趣区域的显着性检测来筛选出视频关键帧。最后,我们通过向传统 LSTM 添加注意力向量来生成视频字幕。我们的方法以 VideoSet 数据集为基准,并通过 BLEU、Meteor 和 Rouge 进行评估。 (C) 2018 Elsevier B.V. 保留所有权利。
Videos captured by people are often tied to certain important moments of their lives. But with the era of big data coming, the time required to retrieval and watch can be daunting. In this paper, novel techniques are proposed for the application of long video segmentation, which can effectively shorten the retrieval time. The motion extent of long video is detected by the improved of the spatio-temporal interest points (STIPs) detection algorithm. After that, the superframe segmentation of the filtered long video is performed to gain the interesting clip of long video. In the selection of keyframes, the region of interest is constructed by the use of the STIP already obtained on the video clips, and the saliency detection of these regions of interest is utilized to screen out video keyframes. Finally, we generate the video captions by adding attention vectors to the traditional LSTM. Our method is benchmarked on the VideoSet dataset, and evaluated by the BLEU, Meteor and Rouge. (C) 2018 Elsevier B.V. All rights reserved.