A long video caption generation algorithm for big video data retrieval
A long video caption generation algorithm for big video data retrieval
复制标题
DOI:
10.1016/j.future.2018.10.054
复制
发表时间:
2019-04-01
影响因子:
7.5
通讯作者:
Wan, Shaohua
中科院分区:
文献类型:
--
作者:
Ding, Songtao;Qu, Shiru;Wan, Shaohua
Videos captured by people are often tied to certain important moments of their lives. But with the era of big data coming, the time required to retrieval and watch can be daunting. In this paper, novel techniques are proposed for the application of long video segmentation, which can effectively shorten the retrieval time. The motion extent of long video is detected by the improved of the spatio-temporal interest points (STIPs) detection algorithm. After that, the superframe segmentation of the filtered long video is performed to gain the interesting clip of long video. In the selection of keyframes, the region of interest is constructed by the use of the STIP already obtained on the video clips, and the saliency detection of these regions of interest is utilized to screen out video keyframes. Finally, we generate the video captions by adding attention vectors to the traditional LSTM. Our method is benchmarked on the VideoSet dataset, and evaluated by the BLEU, Meteor and Rouge. (C) 2018 Elsevier B.V. All rights reserved.