Active Learning from the Web

Active Learning from the Web
复制标题

DOI:
10.1145/3543507.3583346
复制
发表时间:
2022-10
期刊:
Proceedings of the ACM Web Conference 2023
影响因子:
--
通讯作者:
R. Sato
R. Sato
中科院分区:
其他
文献类型:
--
作者:
R. Sato

文献摘要

相似文献

标记数据是机器学习管道中成本最高的过程之一。主动学习是缓解这一问题的标准方法。基于池的主动学习首先建立一个未标记的数据池,然后迭代地选择要标记的数据,从而使所需的标记总数最小化,从而保持模型的高性能。在文献中已经提出了许多从池中选择数据的有效标准。然而,人们对如何建造游泳池的探索较少。具体地说,大多数方法假定特定于任务的池是免费提供的。在本文中,我们主张这样一个特定于任务的池并不总是可用的,并建议使用Web上大量的未标记数据作为应用主动学习的池。由于池非常大,很可能很多任务的池中都存在相关数据,我们不需要为每个任务明确地设计和构建池。挑战在于,由于池的大小,我们不能详尽地计算所有数据的获取分数。本文提出了一种基于用户端信息检索算法的基于主动学习的Web信息检索方法--SEARAING。在实验中,我们使用在线Flickr环境作为主动学习的池。这个资源库包含超过100亿张图片,比现有文献中用于主动学习的资源库大几个数量级。我们确认,我们的方法比现有的使用小的未标记池的方法执行得更好。
Labeling data is one of the most costly processes in machine learning pipelines. Active learning is a standard approach to alleviating this problem. Pool-based active learning first builds a pool of unlabelled data and iteratively selects data to be labeled so that the total number of required labels is minimized, keeping the model performance high. Many effective criteria for choosing data from the pool have been proposed in the literature. However, how to build the pool is less explored. Specifically, most of the methods assume that a task-specific pool is given for free. In this paper, we advocate that such a task-specific pool is not always available and propose the use of a myriad of unlabelled data on the Web for the pool for which active learning is applied. As the pool is extremely large, it is likely that relevant data exist in the pool for many tasks, and we do not need to explicitly design and build the pool for each task. The challenge is that we cannot compute the acquisition scores of all data exhaustively due to the size of the pool. We propose an efficient method, Seafaring, to retrieve informative data in terms of active learning from the Web using a user-side information retrieval algorithm. In the experiments, we use the online Flickr environment as the pool for active learning. This pool contains more than ten billion images and is several orders of magnitude larger than the existing pools in the literature for active learning. We confirm that our method performs better than existing approaches of using a small unlabelled pool.