Lifting Weak Supervision To Structured Prediction

Lifting Weak Supervision To Structured Prediction
复制标题

DOI:
10.48550/arxiv.2211.13375
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Harit Vishwakarma;Nicholas Roberts;Frederic Sala
Harit Vishwakarma;Nicholas Roberts;Frederic Sala
中科院分区:
其他
文献类型:
--
作者:
Harit Vishwakarma;Nicholas Roberts;Frederic Sala

文献摘要

相似文献

弱监督(WS)是一组丰富的技术,通过聚合容易获得的,但潜在的噪声标签估计从各种来源产生伪标签。WS在理论上被很好地理解为二进制分类,其中简单的方法使得能够一致地估计伪标记噪声率。使用这一结果,我们发现在伪标签上训练的下游模型具有与在干净标签上训练的模型几乎相同的泛化保证。虽然这是令人兴奋的,但用户通常希望使用WS进行结构化预测,其中输出空间由不止一个二进制或多类标签集组成:例如排名,图形,流形等。WS对于二元分类的有利理论性质是否提升到这种设置?我们对这个问题的回答是肯定的,在广泛的情况下。对于在有限度量空间中取值的标签,我们引入了基于伪欧几里德嵌入和张量分解的弱监督新技术,提供了一个近乎一致的噪声率估计。对于常曲率黎曼流形中的标签,我们引入了新的不变量,也产生一致的噪声率估计。在这两种情况下,当使用所产生的伪标签与灵活的下游模型相结合时,我们获得的泛化保证几乎与在干净数据上训练的模型相同。我们的几个结果,这可以被视为结构化预测与噪声标签的鲁棒性保证,可能是独立的利益。实证评估验证了我们的主张,并显示所提出的方法的优点。
Weak supervision (WS) is a rich set of techniques that produce pseudolabels by aggregating easily obtained but potentially noisy label estimates from a variety of sources. WS is theoretically well understood for binary classification, where simple approaches enable consistent estimation of pseudolabel noise rates. Using this result, it has been shown that downstream models trained on the pseudolabels have generalization guarantees nearly identical to those trained on clean labels. While this is exciting, users often wish to use WS for structured prediction, where the output space consists of more than a binary or multi-class label set: e.g. rankings, graphs, manifolds, and more. Do the favorable theoretical properties of WS for binary classification lift to this setting? We answer this question in the affirmative for a wide range of scenarios. For labels taking values in a finite metric space, we introduce techniques new to weak supervision based on pseudo-Euclidean embeddings and tensor decompositions, providing a nearly-consistent noise rate estimator. For labels in constant-curvature Riemannian manifolds, we introduce new invariants that also yield consistent noise rate estimation. In both cases, when using the resulting pseudolabels in concert with a flexible downstream model, we obtain generalization guarantees nearly identical to those for models trained on clean data. Several of our results, which can be viewed as robustness guarantees in structured prediction with noisy labels, may be of independent interest. Empirical evaluation validates our claims and shows the merits of the proposed method.