On the role of user-generated metadata in audio visual collections

On the role of user-generated metadata in audio visual collections
复制标题

关于用户生成的元数据在视听收藏中的作用

DOI:
--
复制
发表时间:
2011
期刊:
International Conference on Knowledge Capture
影响因子:
--
通讯作者:
Lora Aroyo
Lora Aroyo
中科院分区:
--
文献类型:
--
作者:
R. Gligorov;M. Hildebrand;J. V. Ossenbruggen;G. Schreiber;Lora Aroyo

文献摘要

被引文献

相似文献

最近,各种众包计划表明,用户社区的有针对性的努力导致了大量的标签。例如,荷兰声音和视觉研究所收集了大量的视频标签游戏Waisda?的标签。为了成功地利用这些标签,需要更好地了解它们的特性。本文的目标是双重的:(i)调查用户在描述视频时使用的词汇,并将其与专业人士使用的词汇进行比较;(ii)确定视频的哪些方面通常被描述,以及使用什么类型的标签。我们报告了对Waisda?收集的标签的分析。关于第一个目标,我们将标签与专业人员使用的典型领域词库以及更通用的词汇表进行了比较。关于第二个目标,我们将标签与视频字幕进行比较,以确定有多少标签来自音频信号。此外,我们进行了定性研究,其中的标签样本解释现有的注释分类框架。结果表明,标签补充专业编目员提供的元数据,标签描述的音频和视频的视觉方面,和用户主要描述对象的视频使用一般的描述。
Recently, various crowdsourcing initiatives showed that targeted efforts of user communities result in massive amounts of tags. For example, the Netherlands Institute for Sound and Vision collected a large number of tags with the video labeling game Waisda?. To successfully utilize these tags, a better understanding of their characteristics is required. The goal of this paper is twofold: (i) to investigate the vocabulary that users employ when describing videos and compare it to the vocabularies used by professionals; and (ii) to establish which aspects of the video are typically described and what type of tags are used for this. We report on an analysis of the tags collected with Waisda?. With respect to the first goal, we compared the the tags with a typical domain thesaurus used by professionals, as well as with a more general vocabulary. With respect to the second goal, we compare the tags to the video subtitles to determine how many tags are derived from the audio signal. In addition, we perform a qualitative study in which a tag sample is interpreted in terms of an existing annotation classification framework. The results suggest that the tags complement the metadata provided by professional cataloguers, the tags describe both the audio and the visual aspects of the video, and the users primarily describe objects in the video using general descriptions.