One-Shot Image Recognition Using Prototypical Encoders with Reduced Hubness

One-Shot Image Recognition Using Prototypical Encoders with Reduced Hubness
复制标题

DOI:
10.1109/wacv48630.2021.00230
复制
发表时间:
2021-01
期刊:
2021 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Chenxi Xiao;Mohammad Norouzi
Chenxi Xiao;Mohammad Norouzi
中科院分区:
其他
文献类型:
--
作者:
Chenxi Xiao;Mohammad Norouzi

文献摘要

相似文献

人类有天生的能力,只需通过观察草图(也称为原型图像)就能识别新物体。类似地,原型图像可以用作看不见的类的有效视觉表示,以解决少镜头学习(FSL)任务。我们的主要目标是识别看不见的手势(手势)交通标志和公司标志,通过他们的图标图像或原型。以前的工作提出利用变分原型编码器(VPE)来解决FSL问题。虽然VPE可以有效地学习图像到图像的翻译任务,但我们发现它的性能受到所谓的中心问题的严重阻碍,并且它无法调节潜在空间中的表示。因此,我们提出了一个新的模型(VPE++),它本质上降低了中心度,并结合了对比和多任务损失,以提高FSL模型的判别能力。结果表明,VPE++的方法可以更好地推广到看不见的类,并可以实现上级的准确性的标志,交通标志和手势数据集相比,国家的最先进的。
Humans have the innate ability to recognize new objects just by looking at sketches of them (also referred as to proto-type images). Similarly, prototypical images can be used as an effective visual representations of unseen classes to tackle few-shot learning (FSL) tasks. Our main goal is to recognize unseen hand signs (gestures) traffic-signs, and corporate-logos, by having their iconographic images or prototypes. Previous works proposed to utilize variational prototypical-encoders (VPE) to address FSL problems. While VPE learns an image-to-image translation task efficiently, we discovered that its performance is significantly hampered by the so-called hubness problem and it fails to regulate the representations in the latent space. Hence, we propose a new model (VPE++) that inherently reduces hubness and incorporates contrastive and multi-task losses to increase the discriminative ability of FSL models. Results show that the VPE++ approach can generalize better to the unseen classes and can achieve superior accuracies on logos, traffic signs, and hand gestures datasets as compared to the state-of-the-art.