Towards Micro-video Understanding by Joint Sequential-Sparse Modeling

Towards Micro-video Understanding by Joint Sequential-Sparse Modeling
复制标题

DOI:
10.1145/3123266.3123341
复制
发表时间:
2017-10
期刊:
Proceedings of the 25th ACM international conference on Multimedia
影响因子:
--
通讯作者:
Meng Liu;Liqiang Nie;M. Wang;Baoquan Chen
Meng Liu;Liqiang Nie;M. Wang;Baoquan Chen
中科院分区:
其他
文献类型:
--
作者:
Meng Liu;Liqiang Nie;M. Wang;Baoquan Chen

文献摘要

被引文献

相似文献

与传统的长视频一样,微视频是文本、声学和视觉形态的统一。这些模态从不同的角度依次讲述现实生活中的事件。然而,与传统的内容丰富的长视频不同,微视频非常短,持续时间为6-15秒,因此它们通常传达一个或几个高层次的概念。鉴于此,我们必须对稀疏性和多个序列结构进行表征和联合建模,以更好地理解微视频。为了实现这一点,在本文中,我们提出了一个端到端的深度学习模型,该模型包含三个并行的LSTM来捕获序列结构,以及一个卷积神经网络来学习微视频的稀疏概念级表示。将该模型应用于微视频分类。此外,我们还构建了一个真实世界的序列建模数据集,并将其发布,以方便其他研究人员。实验结果表明,我们的模型产生更好的性能比几个国家的最先进的基线。
Like the traditional long videos, micro-videos are the unity of textual, acoustic, and visual modalities. These modalities sequentially tell a real-life event from distinct angles. Yet, unlike the traditional long videos with rich content, micro-videos are very short, lasting for 6-15 seconds, and they hence usually convey one or a few high-level concepts. In the light of this, we have to characterize and jointly model the sparseness and multiple sequential structures for better micro-video understanding. To accomplish this, in this paper, we present an end-to-end deep learning model, which packs three parallel LSTMs to capture the sequential structures and a convolutional neural network to learn the sparse concept-level representations of micro-videos. We applied our model to the application of micro-video categorization. Besides, we constructed a real-world dataset for sequence modeling and released it to facilitate other researchers. Experimental results demonstrate that our model yields better performance than several state-of-the-art baselines.