Cross-modal correlation learning for clustering on image-audio dataset

Cross-modal correlation learning for clustering on image-audio dataset
复制标题

DOI:
10.1145/1291233.1291290
复制
发表时间:
2007-09
期刊:
Proceedings of the 15th ACM international conference on Multimedia
影响因子:
--
通讯作者:
Hong Zhang;Yueting Zhuang;Fei Wu
Hong Zhang;Yueting Zhuang;Fei Wu
中科院分区:
其他
文献类型:
--
作者:
Hong Zhang;Yueting Zhuang;Fei Wu

文献摘要

被引文献

相似文献

探索不同数据集之间的相关性并利用这些相关性对这些数据集进行聚类是有趣且具有挑战性的。图像和音频之间的跨模态相关性可以帮助识别某些语义的图像(或音频)。然而,异质性问题使得学习视觉和听觉特征之间的跨模态相关性变得困难。本文首先分析了子空间映射过程中图像和音频特征矩阵之间的典型相关性,然后设计了基于相关性的图像和音频相似性增强算法,最后利用仿射传播算法实现了图像聚类和音频聚类。在图像-音频数据集上的实验结果令人鼓舞,表明我们的方法是有效的。我们给出了一个有趣的应用程序查询图像的音频的例子。
It is interesting and challenging to explore correlations between different datasets and utilize such correlations for the clustering on these datasets. Cross-modal correlation between images and audios can help identify images (or audios) of certain semantics. However, the heterogeneous problem makes it difficult to learn cross-modal correlation between visual and auditory features. In this paper, we analyze canonical correlation between feature matrices of images and audios during subspace mapping; then we design correlation-based similarity reinforcement for images and audios; thirdly we implement image clustering and audio clustering with affinity propagation. Experiment results on image-audio dataset are encouraging and show that the performance of our approach is effective. We give an interesting application of querying images by audio examples.