Universalizing Weak Supervision

Universalizing Weak Supervision
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Changho Shin;Winfred Li;Harit Vishwakarma;Nicholas Roberts;Frederic Sala
Changho Shin;Winfred Li;Harit Vishwakarma;Nicholas Roberts;Frederic Sala
中科院分区:
其他
文献类型:
--
作者:
Changho Shin;Winfred Li;Harit Vishwakarma;Nicholas Roberts;Frederic Sala

文献摘要

相似文献

弱监督(WS)框架是绕过手工标记的大型数据集的一种流行方式,用于培训数据渴望的模型。这些方法合成了多个嘈杂的标签,但廉价获得的标签估计值是一组高质量的伪标记,用于下游训练。但是,合成技术是特定于特定类型的标签(例如二进制标签或序列)的,并且每种新标签类型都需要手动设计新的合成算法。取而代之的是,我们提出了一种通用技术,该技术能够对任何标签类型进行弱监督,同时仍提供理想的属性,包括实际的灵活性,计算效率和理论保证。我们将此技术应用于以前没有解决WS框架在内的重要问题,包括在双曲线空间中学习排名,回归和学习。从理论上讲,我们的合成方法产生了一个一致的估计量,以学习指数家庭模型的一些具有挑战性但重要的概括。在实验上,我们验证了我们的框架,并在包括现实世界中的学习到级别和回归问题(以及对双曲线歧管的学习)中的基准中对基准的改进。
Weak supervision (WS) frameworks are a popular way to bypass hand-labeling large datasets for training data-hungry models. These approaches synthesize multiple noisy but cheaply-acquired estimates of labels into a set of high-quality pseudolabels for downstream training. However, the synthesis technique is specific to a particular kind of label, such as binary labels or sequences, and each new label type requires manually designing a new synthesis algorithm. Instead, we propose a universal technique that enables weak supervision over any label type while still offering desirable properties, including practical flexibility, computational efficiency, and theoretical guarantees. We apply this technique to important problems previously not tackled by WS frameworks including learning to rank, regression, and learning in hyperbolic space. Theoretically, our synthesis approach produces a consistent estimators for learning some challenging but important generalizations of the exponential family model. Experimentally, we validate our framework and show improvement over baselines in diverse settings including real-world learning-to-rank and regression problems along with learning on hyperbolic manifolds.