Data Programming: Creating Large Training Sets, Quickly

Data Programming: Creating Large Training Sets, Quickly
复制标题

DOI:
--
复制
发表时间:
2016-05
期刊:
Advances in neural information processing systems
影响因子:
--
通讯作者:
Alexander J. Ratner;Christopher De Sa;Sen Wu;Daniel Selsam;C. Ré
Alexander J. Ratner;Christopher De Sa;Sen Wu;Daniel Selsam;C. Ré
中科院分区:
其他
文献类型:
--
作者:
Alexander J. Ratner;Christopher De Sa;Sen Wu;Daniel Selsam;C. Ré

文献摘要

被引文献

相似文献

大型标记训练集是监督学习方法的关键构建块,也是深度学习技术的关键推动因素。对于某些应用程序,创建标记的训练集是应用机器学习中最耗时和昂贵的部分。因此,我们提出了一个范例的编程创建的训练集称为数据编程,其中用户表示弱监督策略或域的标记功能,这是程序的标签子集的数据,但嘈杂,可能会发生冲突。我们表明,通过明确地将此训练集标记过程表示为生成模型,我们可以对生成的训练集进行“降噪”,并从理论上建立我们可以在少数设置中恢复这些生成模型的参数。然后,我们将展示如何修改判别损失函数,使其具有噪声意识,并在一系列判别模型(包括逻辑回归和LSTM)上演示我们的方法。在2014年的TAC-KBP Slot Filling挑战赛中,我们通过实验证明了数据编程会带来一个新的获胜分数,并且还证明了将数据编程应用于LSTM模型会使TAC-KBP的分数比最先进的LSTM基线高出近6个F1点(并在比赛中获得第二名)。此外,在最初的用户研究中,我们观察到,当训练数据有限或不可用时,数据编程可能是非专家创建机器学习模型的更简单方法。
Large labeled training sets are the critical building blocks of supervised learning methods and are key enablers of deep learning techniques. For some applications, creating labeled training sets is the most time-consuming and expensive part of applying machine learning. We therefore propose a paradigm for the programmatic creation of training sets called data programming in which users express weak supervision strategies or domain heuristics as labeling functions, which are programs that label subsets of the data, but that are noisy and may conflict. We show that by explicitly representing this training set labeling process as a generative model, we can "denoise" the generated training set, and establish theoretically that we can recover the parameters of these generative models in a handful of settings. We then show how to modify a discriminative loss function to make it noise-aware, and demonstrate our method over a range of discriminative models including logistic regression and LSTMs. Experimentally, on the 2014 TAC-KBP Slot Filling challenge, we show that data programming would have led to a new winning score, and also show that applying data programming to an LSTM model leads to a TAC-KBP score almost 6 F1 points over a state-of-the-art LSTM baseline (and into second place in the competition). Additionally, in initial user studies we observed that data programming may be an easier way for non-experts to create machine learning models when training data is limited or unavailable.