Simultaneous Realization of Multiple Music Video Applications Based on Heterogeneous Network Analysis Via Latent Link Estimation

Simultaneous Realization of Multiple Music Video Applications Based on Heterogeneous Network Analysis Via Latent Link Estimation
复制标题

DOI:
10.1109/icme.2018.8486474
复制
发表时间:
2018-07
期刊:
2018 IEEE International Conference on Multimedia and Expo (ICME)
影响因子:
--
通讯作者:
Yui Matsumoto;Ryosuke Harakawa;Takahiro Ogawa;M. Haseyama
Yui Matsumoto;Ryosuke Harakawa;Takahiro Ogawa;M. Haseyama
中科院分区:
其他
文献类型:
--
作者:
Yui Matsumoto;Ryosuke Harakawa;Takahiro Ogawa;M. Haseyama

文献摘要

相似文献

为了帮助用户寻找所需的音乐视频和创建有吸引力的音乐视频,已经提出了许多实现诸如音乐视频推荐、字幕和生成等应用的方法。在本文中,同时实现这些应用程序的异构网络分析的基础上,通过潜在链接估计的新方法,提出了。据我们所知,这项工作是第一次尝试,同时实现音乐视频推荐,字幕和生成。所提出的方法,使潜在的链接估计考虑多模态信息和多个社会元数据从音乐视频通过拉普拉斯多集典型相关分析。因此,它成为可行的,以构建一个异构网络,使音频,视频和文本信息的音乐视频和用户信息在同一特征空间上的直接比较。此外,对所获得的异构网络的链接预测使得能够与(i)用户信息和他们期望的音频信息;(ii)描述乐曲的内容的音频信息和文本信息;以及(iii)可视地表示乐曲的内容的音频信息和可视信息相关联。因此,对(i)音乐视频推荐的支持;(ii)字幕;和(iii)生成分别变得可行。利用YouTube-8 M构建的真实世界数据集的实验结果表明了该方法的有效性。
To help users seek desired music videos and create attractive music videos, many methods that realize applications such as music video recommendation, captioning and generation have been proposed. In this paper, a novel method that realizes these applications simultaneously on the basis of heterogeneous network analysis via latent link estimation is proposed. To the best of our knowledge, this work is the first attempt to realize music video recommendation, captioning and generation simultaneously. The proposed method enables latent link estimation with consideration of multimodal information and multiple social metadata obtained from music videos via Laplacian multiset canonical correlation analysis. Thus, it becomes feasible to construct a heterogeneous network that enables direct comparison of audio, visual and textual information of music videos and user information on the same feature space. Furthermore, link prediction on the obtained heterogeneous network enables association with (i) user information and their desired audio information; (ii) audio information and textual information that describes contents of musical pieces; and (iii) audio information and visual information that represents contents of musical pieces visually. As a result, support for (i) music video recommendation; (ii) captioning; and (iii) generation becomes feasible, respectively. Experimental results for a real-world dataset constructed by using YouTube-8M show the effectiveness of the proposed method.