Multiple Bernoulli relevance models for image and video annotation

Multiple Bernoulli relevance models for image and video annotation
复制标题

DOI:
10.1109/cvpr.2004.171
复制
发表时间:
2004-06
期刊:
Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004.
影响因子:
--
通讯作者:
Shaolei Feng;R. Manmatha;V. Lavrenko
Shaolei Feng;R. Manmatha;V. Lavrenko
中科院分区:
其他
文献类型:
--
作者:
Shaolei Feng;R. Manmatha;V. Lavrenko

文献摘要

被引文献

相似文献

响应于文本查询检索图像需要对图片的语义有一定的了解。在这里,我们展示了如何使用多重伯努利相关性模型从图像和视频中进行自动图像标注和检索(使用一个单词查询)。该模型假设提供了带有关键字注释的图像或视频的训练集。为图像提供多个关键字,并且不提供关键字和图像之间的特定对应关系。每幅图像被分割成一组矩形区域,并在这些区域上计算实值特征向量。相关性模型是单词标注和图像特征向量的联合概率分布,并使用训练集来计算。使用多重伯努利模型估计单词概率,并使用非参数核密度估计来估计图像特征概率。然后使用该模型对测试集中的图像进行注释。我们在标准Corel数据集的图像和NIST视频树的一组视频关键帧上进行了实验。对比实验表明,该模型比基于流行的多项分布估计单词概率的模型具有更好的性能。实验结果还表明,我们的模型在图像和视频标注任务上的性能明显优于已有的研究结果。
Retrieving images in response to textual queries requires some knowledge of the semantics of the picture. Here, we show how we can do both automatic image annotation and retrieval (using one word queries) from images and videos using a multiple Bernoulli relevance model. The model assumes that a training set of images or videos along with keyword annotations is provided. Multiple keywords are provided for an image and the specific correspondence between a keyword and an image is not provided. Each image is partitioned into a set of rectangular regions and a real-valued feature vector is computed over these regions. The relevance model is a joint probability distribution of the word annotations and the image feature vectors and is computed using the training set. The word probabilities are estimated using a multiple Bernoulli model and the image feature probabilities using a non-parametric kernel density estimate. The model is then used to annotate images in a test set. We show experiments on both images from a standard Corel data set and a set of video key frames from NIST's video tree. Comparative experiments show that the model performs better than a model based on estimating word probabilities using the popular multinomial distribution. The results also show that our model significantly outperforms previously reported results on the task of image and video annotation.