Revisiting metric learning for few -shot image classification

Revisiting metric learning for few -shot image classification
复制标题

DOI:
10.1016/j.neucom.2020.04.040
复制
发表时间:
2020-09-17
期刊:
影响因子:
6
通讯作者:
Heng, Pheng-Ann
Heng, Pheng-Ann
中科院分区:
计算机科学2区
文献类型:
--
作者:
Li, Xiaomeng;Yu, Lequan;Heng, Pheng-Ann

文献摘要

被引文献

相似文献

少镜头学习的目标是在每个类中仅使用少量标记样本来识别新的视觉概念。最近有效的基于度量的少数镜头方法采用神经网络来学习查询和支持示例之间的特征相似性比较。然而,特征嵌入的重要性,即,训练样本之间的关系,被忽略了。在这项工作中,我们提出了一个简单而强大的基线,通过强调特征嵌入的重要性,少拍分类。具体来说,我们重新审视了深度度量学习中的经典三元组网络,并将其扩展为用于少量学习的deepK-tuplet网络,利用输入样本之间的关系通过episode训练来学习一般表示学习。一旦经过训练,我们的网络就能够为看不见的新类别提取有区别的特征,并且可以与非线性距离度量函数无缝结合,以促进少数分类。我们在miniImageNet基准测试中的结果优于其他基于度量的少数分类方法。更重要的是,当使用miniImageNet训练的模型在完全不同的数据集(加州理工学院-101,CUB-200,斯坦福大学狗和汽车)上进行评估时,我们的方法显着优于以前的方法,证明了其上级泛化能力。
The goal of few-shot learning is to recognize new visual concepts with just a few amount of labeled samples in each class. Recent effective metric-based few-shot approaches employ neural networks to learn a feature similarity comparison between query and support examples. However, the importance of feature embedding,i.e., exploring the relationship among training samples, is neglected. In this work, we present a simple yet powerful baseline for few-shot classification by emphasizing the importance of feature embedding. Specifically, we revisit the classical triplet network from deep metric learning, and extend it into a deepK-tuplet network for few-shot learning, utilizing the relationship among the input samples to learn a general representation learning via episode-training. Once trained, our network is able to extract discriminative features for unseen novel categories and can be seamlessly incorporated with a non-linear distance metric function to facilitate the few-shot classification. Our result on the miniImageNet benchmark outperforms other metric-based few-shot classification methods. More importantly, when evaluated on completely different datasets (Caltech-101, CUB-200, Stanford Dogs and Cars) using the model trained with miniImageNet, our method significantly outperforms prior methods, demonstrating its superior capability to generalize to unseen classes.