Active metric learning and classification using similarity queries

Active metric learning and classification using similarity queries
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
--
影响因子:
--
通讯作者:
Namrata Nadagouda;Austin Xu;M. Davenport
Namrata Nadagouda;Austin Xu;M. Davenport
中科院分区:
其他
文献类型:
--
作者:
Namrata Nadagouda;Austin Xu;M. Davenport

文献摘要

相似文献

主动学习通常用于通过自适应地选择信息量最大的查询来训练标签效率模型。然而,大多数主动学习策略被设计为学习数据的表示(例如,嵌入或度量学习)或在任务上表现良好(例如,分类)的数据。然而,许多机器学习任务涉及表征学习和特定任务目标的组合。出于这一动机,我们提出了一种新的统一查询框架,可以应用于任何问题,其中的一个关键组成部分是学习的数据表示,反映相似性。我们的方法建立在相似性或最近邻(NN)查询的基础上,这些查询旨在选择导致改进嵌入的样本。查询由引用和一组对象组成,其中oracle选择最相似的对象(即,最接近)的参考。为了减少请求查询的数量,根据信息理论标准自适应地选择它们。我们证明了所提出的策略在两个任务上的有效性-主动度量学习和主动分类-使用各种合成和真实的世界数据集。特别是,我们证明了在深度度量学习环境中,主动选择的NN查询优于最近开发的主动三元组选择方法。此外,我们表明,在分类,积极选择类标签可以重新制定为一个过程,选择最翔实的NN查询,允许直接应用我们的方法。
Active learning is commonly used to train label-efficient models by adaptively selecting the most informative queries. However, most active learning strategies are designed to either learn a representation of the data (e.g., embedding or metric learning) or perform well on a task (e.g., classification) on the data. However, many machine learning tasks involve a combination of both representation learning and a task-specific goal. Motivated by this, we propose a novel unified query framework that can be applied to any problem in which a key component is learning a representation of the data that reflects similarity. Our approach builds on similarity or nearest neighbor (NN) queries which seek to select samples that result in improved embeddings. The queries consist of a reference and a set of objects, with an oracle selecting the object most similar (i.e., nearest) to the reference. In order to reduce the number of solicited queries, they are chosen adaptively according to an information theoretic criterion. We demonstrate the effectiveness of the proposed strategy on two tasks -- active metric learning and active classification -- using a variety of synthetic and real world datasets. In particular, we demonstrate that actively selected NN queries outperform recently developed active triplet selection methods in a deep metric learning setting. Further, we show that in classification, actively selecting class labels can be reformulated as a process of selecting the most informative NN query, allowing direct application of our method.