City-Identification of Flickr Videos Using Semantic Acoustic Features

City-Identification of Flickr Videos Using Semantic Acoustic Features
复制标题

使用语义声学特征进行 Flickr 视频的城市识别

DOI:
--
复制
发表时间:
2016
期刊:
IEEE International Conference on Multimedia Big Data
影响因子:
--
通讯作者:
Ian Lane
Ian Lane
中科院分区:
--
文献类型:
--
作者:
Benjamin Elizalde;Guan;Mingzhi Zeng;Ian Lane

文献摘要

参考文献

被引文献

相似文献

视频的城市识别旨在确定视频属于一组城市的可能性。在本文中,我们提出了一种只使用音频的方法,因此我们不使用任何额外的形式,如图像,用户标签或地理标签。通过这种方式,我们展示了视频的城市位置与其声学信息相关的程度。这项任务的成功表明,可以作出改进,以补充其他模式。特别是,我们提出了一种方法来计算和使用语义声学特征来执行城市识别和特征显示的语义证据的识别。语义的证据是由城市声音的分类,并表示这些声音在城市音轨的潜在存在。我们使用MediaEval Placing Task集合,其中包含按城市标记的Flickr视频。此外,我们使用了UrbanSound8K集,其中包含按声音类型标记的音频片段。我们的方法提高了最先进的性能,并提供了一种新的语义方法来完成这项任务。
City-identification of videos aims to determine the likelihood of a video belonging to a set of cities. In this paper, we present an approach using only audio, thus we do not use any additional modality such as images, user-tags or geo-tags. In this manner, we show to what extent the city-location of videos correlates to their acoustic information. Success in this task suggests improvements can be made to complement the other modalities. In particular, we present a method to compute and use semantic acoustic features to perform city-identification and the features show semantic evidence of the identification. The semantic evidence is given by a taxonomy of urban sounds and expresses the potential presence of these sounds in the city-soundtracks. We used the MediaEval Placing Task set, which contains Flickr videos labeled by city. In addition, we used the UrbanSound8K set containing audio clips labeled by sound-type. Our method improved the state-of-the-art performance and provides a novel semantic approach to this task.
DOI: 10.1109/msp.2014.2326181
发表时间: 2015-05-01
影响因子: 14.9
作者:
Barchiesi, Daniele;Giannoulis, Dimitrios;Plumbley, Mark D.
通讯作者: Plumbley, Mark D.