AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels

AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels
复制标题

DOI:
10.48550/arxiv.2208.14362
复制
发表时间:
2022-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Nicholas Roberts;Xintong Li;Tzu-Heng Huang;Dyah Adila;Spencer Schoenberg;Chengao Liu;Lauren Pick;Haotian Ma;Aws Albarghouthi;Frederic Sala
Nicholas Roberts;Xintong Li;Tzu-Heng Huang;Dyah Adila;Spencer Schoenberg;Chengao Liu;Lauren Pick;Haotian Ma;Aws Albarghouthi;Frederic Sala
中科院分区:
其他
文献类型:
--
作者:
Nicholas Roberts;Xintong Li;Tzu-Heng Huang;Dyah Adila;Spencer Schoenberg;Chengao Liu;Lauren Pick;Haotian Ma;Aws Albarghouthi;Frederic Sala

文献摘要

被引文献

相似文献

弱监督(WS)是一种强大的方法,用于在很少或没有标记数据的情况下构建标记数据集来训练监督模型。它将手工标记数据替换为由标记函数(LFs)表示的聚合多个噪声但便宜的标记估计。虽然弱监督在许多领域得到了成功的应用,但对于具有复杂或高维特征的领域,弱监督难以构建标记函数,限制了其应用范围。为了解决这个问题,已经提出了一些方法,使用一组少量的地面真值标签来自动化LF设计过程。在这项工作中,我们介绍了AutoWS- bench -101:一个在具有挑战性的WS设置中评估自动化WS (AutoWS)技术的框架——一组不同的应用领域,在这些领域上以前很难或不可能应用传统的WS技术。虽然AutoWS是扩大WS应用范围的一个有希望的方向,但零射击基础模型等强大方法的出现表明,需要了解AutoWS技术如何与现代零射击或少射击学习器进行比较或合作。这告知了AutoWS- bench -101的中心问题:给定每个任务的100个标签的初始集,我们询问从业者是否应该使用AutoWS方法来生成额外的标签,或者使用一些更简单的基线,例如来自基础模型或监督学习的零概率预测。我们观察到,在许多情况下,如果要优于简单的少量基线,AutoWS方法有必要纳入来自基础模型的信号,而AutoWS- bench -101促进了这一方向的未来研究。最后,我们对AutoWS方法进行了彻底的消融研究。
Weak supervision (WS) is a powerful method to build labeled datasets for training supervised models in the face of little-to-no labeled data. It replaces hand-labeling data with aggregating multiple noisy-but-cheap label estimates expressed by labeling functions (LFs). While it has been used successfully in many domains, weak supervision's application scope is limited by the difficulty of constructing labeling functions for domains with complex or high-dimensional features. To address this, a handful of methods have proposed automating the LF design process using a small set of ground truth labels. In this work, we introduce AutoWS-Bench-101: a framework for evaluating automated WS (AutoWS) techniques in challenging WS settings -- a set of diverse application domains on which it has been previously difficult or impossible to apply traditional WS techniques. While AutoWS is a promising direction toward expanding the application-scope of WS, the emergence of powerful methods such as zero-shot foundation models reveals the need to understand how AutoWS techniques compare or cooperate with modern zero-shot or few-shot learners. This informs the central question of AutoWS-Bench-101: given an initial set of 100 labels for each task, we ask whether a practitioner should use an AutoWS method to generate additional labels or use some simpler baseline, such as zero-shot predictions from a foundation model or supervised learning. We observe that in many settings, it is necessary for AutoWS methods to incorporate signal from foundation models if they are to outperform simple few-shot baselines, and AutoWS-Bench-101 promotes future research in this direction. We conclude with a thorough ablation study of AutoWS methods.