Compressing tags to find interesting media groups

Compressing tags to find interesting media groups
复制标题

压缩标签以查找有趣的媒体组

DOI:
--
复制
发表时间:
2009
期刊:
International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
A. Siebes
A. Siebes
中科院分区:
--
文献类型:
--
作者:
M. Leeuwen;F. Bonchi;Börkur Sigurbjörnsson;A. Siebes

文献摘要

被引文献

相似文献

在像Flickr和Zooomr这样的照片分享网站上,用户可以为他们上传的照片分配标签。使用这些标记在给定查询的结果集中查找语义相关的有趣图片组是一个明显的应用程序问题。我们从最小描述长度(MDL)的角度分析了这个问题,并开发了一个算法来找到最有趣的组。该方法基于Krimp,该方法通过压缩找到具有数据特征的小组模式。这些模式是一组标签,通常一起分配给照片。数据库压缩得越好,它包含的结构就越多,因此就越同质。根据这一观察,我们设计了一个基于压缩的测量方法。我们在Flickr数据上的实验表明,我们找到了最有趣、最同质的群体。我们展示了大量的示例,并与Flickr网站上的聚类进行了比较。
On photo sharing websites like Flickr and Zooomr, users are offered the possibility to assign tags to their uploaded pictures. Using these tags to find interesting groups of semantically related pictures in the result set of a given query is a problem with obvious applications. We analyse this problem from a Minimum Description Length (MDL) perspective and develop an algorithm that finds the most interesting groups. The method is based on Krimp, which finds small sets of patterns that characterise the data using compression. These patterns are sets of tags, often assignedtogether to photos. The better a database compresses, the more structure it contains and thus the more homogeneous it is. Following this observation we devise a compression-based measure. Our experiments on Flickr data show that the most interesting and homogeneous groups are found. We show extensive examples and compare to clusterings on the Flickr website.