Generative Modeling Helps Weak Supervision (and Vice Versa)

Generative Modeling Helps Weak Supervision (and Vice Versa)
复制标题

DOI:
10.48550/arxiv.2203.12023
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Benedikt Boecking;W. Neiswanger;Nicholas Roberts;Stefano Ermon;Frederic Sala;A. Dubrawski
Benedikt Boecking;W. Neiswanger;Nicholas Roberts;Stefano Ermon;Frederic Sala;A. Dubrawski
中科院分区:
其他
文献类型:
--
作者:
Benedikt Boecking;W. Neiswanger;Nicholas Roberts;Stefano Ermon;Frederic Sala;A. Dubrawski

文献摘要

相似文献

监督机器学习的许多有前途的应用在获取足够数量和质量的标记数据方面面临障碍,这造成了昂贵的瓶颈。为了克服这些限制,人们研究了不依赖于地面真实标签的技术,包括弱监督和生成式建模。虽然这些技术似乎可以协同使用,相互改进,但如何在它们之间构建接口还没有得到很好的理解。在这项工作中,我们提出了一个融合程序化弱监督和生成对抗网络的模型,并提供了激励这种融合的理论依据。所提出的方法捕获数据中的离散潜变量以及弱监督导出的标签估计。两者的对齐允许更好地建模弱监督源的样本相关精度,提高未观察到的标签的估计。这是第一种通过弱监督合成图像和伪标签实现数据增强的方法。此外,其学习的潜变量可以定性地检查。该模型在许多多类图像分类数据集上优于基线弱监督标签模型,提高了生成图像的质量,并通过合成样本的数据增强进一步提高了最终模型的性能。
Many promising applications of supervised machine learning face hurdles in the acquisition of labeled data in sufficient quantity and quality, creating an expensive bottleneck. To overcome such limitations, techniques that do not depend on ground truth labels have been studied, including weak supervision and generative modeling. While these techniques would seem to be usable in concert, improving one another, how to build an interface between them is not well-understood. In this work, we propose a model fusing programmatic weak supervision and generative adversarial networks and provide theoretical justification motivating this fusion. The proposed approach captures discrete latent variables in the data alongside the weak supervision derived label estimate. Alignment of the two allows for better modeling of sample-dependent accuracies of the weak supervision sources, improving the estimate of unobserved labels. It is the first approach to enable data augmentation through weakly supervised synthetic images and pseudolabels. Additionally, its learned latent variables can be inspected qualitatively. The model outperforms baseline weak supervision label models on a number of multiclass image classification datasets, improves the quality of generated images, and further improves end-model performance through data augmentation with synthetic samples.