Discover-Then-Rank Unlabeled Support Vectors in the Dual Space for Multi-Class Active Learning

Discover-Then-Rank Unlabeled Support Vectors in the Dual Space for Multi-Class Active Learning
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Dayou Yu;Weishi Shi;Qi Yu
Dayou Yu;Weishi Shi;Qi Yu
中科院分区:
其他
文献类型:
--
作者:
Dayou Yu;Weishi Shi;Qi Yu

文献摘要

相似文献

通过利用稀疏核最大间隔预测器对偶空间的关键性质,我们提出了一种从发现潜在支持向量并对其进行排序的新视角来实现主动学习。我们从理论上分析了铰链损耗在对偶形式下的变化,给出了与对偶空间的关键几何性质密切相关的上下界,从而帮助我们识别出AL的各种重要数据样本。这些界限有助于设计一种新的抽样策略,该策略利用分类证据作为关键工具,通过对偶变量和核评估的仿射组合形成。我们构造了两种不同类型的抽样函数,包括1)发现,它关注所有类别中总证据较少的样本,以支持探索;2)排名,目的是进一步细化决策边界。这两个功能被自动安排到两个阶段的主动采样过程中,以平衡勘探和开发。在各种真实数据上的实验表明,我们的模型达到了最先进的AL性能。
We propose to approach active learning (AL) from a novel perspective of discovering and then ranking potential support vectors by leveraging the key properties of the dual space of a sparse kernel max-margin predictor. We theoretically analyze the change of a hinge loss in the dual form and provide both the upper and lower bounds that are deeply connected to the key geometric properties induced by the dual space, which then help us identify various types of important data samples for AL. These bounds inform the design of a novel sampling strategy that leverages class-wise evidence as a key vehicle, formed through an affine combination of dual variables and kernel evaluation. We construct two distinct types of sampling functions, including 1) discovery, which focuses on samples with low total evidence from all classes to support exploration, and 2) ranking, which aims to further refine the decision boundary. These two functions are automatically arranged into a two-phase active sampling process to balance exploration and exploitation. Experiments on various real-world data demonstrate the state-of-the-art AL performance achieved by our model.