Deep Metric Learning for Multi-Label and Multi-Object Image Retrieval

Deep Metric Learning for Multi-Label and Multi-Object Image Retrieval
复制标题

DOI:
10.1587/transinf.2020edp7226
复制
发表时间:
2021-06
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Jonathan Mojoo;Takio Kurita
Jonathan Mojoo;Takio Kurita
中科院分区:
其他
文献类型:
--
作者:
Jonathan Mojoo;Takio Kurita

文献摘要

相似文献

长期以来,基于内容的图像检索一直是计算机视觉研究者的研究热点。多年来,深度度量学习取得了许多进展,最近的一个进展是深度度量学习,其灵感来自于深度神经网络在许多机器学习任务中的成功。度量学习的目标是利用神经网络从图像像素数据中提取好的高级特征。这些特征提供了有用的抽象,使算法能够以类似人类的精度在图像之间进行视觉比较。为了学习这些特征,通常使用图像相似度或相对相似度的监督信息。深度度量学习的一个重要问题是如何定义图像中多标签或多目标场景的相似性。传统上,两两相似性是基于两个图像之间存在单个公共标签来定义的。然而,这个定义非常粗糙,不适合多标签或多对象数据。另一个常见的错误是完全忽略图像中对象的多样性,从而忽略了某些类型数据集的多对象方面。在我们的工作中,我们提出了一种基于多标签和多目标图像数据的相对相似性来学习深度图像表示的方法。我们引入了一种基于Jaccard相似性系数的直观有效的相似性度量,它等价于两个标签集的交集/并集。因此,我们将相似性视为连续的,而不是离散的量。我们将这种相似性度量纳入具有自适应余量的三重损失中,并在图像检索任务中获得了良好的平均精度。我们进一步证明,使用最近提出的量化方法,得到的深度特征可以量化,同时保持相似性。我们还表明,我们提出的相似度度量比以前提出的基于余弦相似度的度量在多目标图像上表现得更好。我们提出的方法在两个基准数据集上优于几种最先进的方法。
SUMMARY Content-based image retrieval has been a hot topic among computer vision researchers for a long time. There have been many advances over the years, one of the recent ones being deep metric learning, inspired by the success of deep neural networks in many machine learning tasks. The goal of metric learning is to extract good high-level features from image pixel data using neural networks. These features provide useful abstractions, which can enable algorithms to perform visual comparison between images with human-like accuracy. To learn these features, supervised information of image similarity or relative similarity is often used. One important issue in deep metric learning is how to define similarity for multi-label or multi-object scenes in images. Traditionally, pairwise similarity is defined based on the presence of a single common label between two images. However, this definition is very coarse and not suitable for multi-label or multi-object data. Another common mistake is to completely ignore the multiplicity of objects in images, hence ignoring the multi-object facet of certain types of datasets. In our work, we propose an approach for learning deep image representations based on the relative similarity of both multi-label and multi-object image data. We introduce an intuitive and effective similarity metric based on the Jaccard similarity coe ffi cient, which is equivalent to the intersection over union of two label sets. Hence we treat similarity as a continuous, as opposed to discrete quantity. We incorporate this similarity metric into a triplet loss with an adaptive margin, and achieve good mean average precision on image retrieval tasks. We further show, using a recently proposed quantization method, that the resulting deep feature can be quantized whilst preserving similarity. We also show that our proposed similarity metric performs better for multi-object images than a previously proposed cosine similarity-based metric. Our proposed method outperforms several state-of-the-art methods on two benchmark datasets.