SIMILAR: Submodular Information Measures Based Active Learning In Realistic Scenarios

SIMILAR: Submodular Information Measures Based Active Learning In Realistic Scenarios
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
S. Kothawade;Nathan Beck;Krishnateja Killamsetty;Rishabh K. Iyer
S. Kothawade;Nathan Beck;Krishnateja Killamsetty;Rishabh K. Iyer
中科院分区:
其他
文献类型:
--
作者:
S. Kothawade;Nathan Beck;Krishnateja Killamsetty;Rishabh K. Iyer

文献摘要

相似文献

主动学习已被证明是有用的,通过选择信息量最大的样本,最大限度地减少标签成本。然而,现有的主动学习方法在现实场景中不能很好地工作,例如不平衡或稀有类,未标记集中的分布数据和冗余。在这项工作中,我们提出了SIMILAR(基于子模块信息测量的主动学习),这是一个统一的主动学习框架,使用最近提出的子模块信息测量(SIM)作为获取函数。我们认为,SIMILAR不仅适用于标准的主动学习,而且还可以轻松扩展到上面考虑的现实环境,并作为一个一站式的主动学习解决方案,可扩展到大型现实世界的数据集。从经验上讲,我们表明,SIMILAR在稀有类的情况下比现有的主动学习算法性能高出约5% - 18%,在CIFAR-10,MNIST和ImageNet等几个图像分类任务的分布数据的情况下比现有的主动学习算法性能高出约5% - 10%。SIMILAR是DISTIL工具包的一部分:“https://github.com/decile-team/distil“。
Active learning has proven to be useful for minimizing labeling costs by selecting the most informative samples. However, existing active learning methods do not work well in realistic scenarios such as imbalance or rare classes, out-of-distribution data in the unlabeled set, and redundancy. In this work, we propose SIMILAR (Submodular Information Measures based actIve LeARning), a unified active learning framework using recently proposed submodular information measures (SIM) as acquisition functions. We argue that SIMILAR not only works in standard active learning, but also easily extends to the realistic settings considered above and acts as a one-stop solution for active learning that is scalable to large real-world datasets. Empirically, we show that SIMILAR significantly outperforms existing active learning algorithms by as much as ~5% - 18% in the case of rare classes and ~5% - 10% in the case of out-of-distribution data on several image classification tasks like CIFAR-10, MNIST, and ImageNet. SIMILAR is available as a part of the DISTIL toolkit:"https://github.com/decile-team/distil".