Snorkel: Rapid Training Data Creation with Weak Supervision.

Snorkel: Rapid Training Data Creation with Weak Supervision.
复制标题

DOI:
10.14778/3157794.3157797
复制
发表时间:
2017-11
期刊:
Proceedings of the VLDB Endowment. International Conference on Very Large Data Bases
影响因子:
--
通讯作者:
Ré C
Ré C
中科院分区:
其他
文献类型:
--
作者:
Ratner A;Bach SH;Ehrenberg H;Fries J;Wu S;Ré C

文献摘要

参考文献

被引文献

相似文献

标注训练数据日益成为部署机器学习系统的最大瓶颈。我们推出了Snorkel,这是一种首个此类系统,使用户能够训练最先进的模型,而无需手动标记任何训练数据。取而代之的是,用户编写表达任意启发式的标签函数,其精度和相关性可能未知。Snorkel通过整合我们最近提出的机器学习范式的第一个端到端实现-数据编程-来消除他们的输出,而不需要获得基本的事实。我们根据过去一年与公司、机构和研究实验室合作的经验,为编写标签函数提供了一个灵活的接口层。在一项用户研究中,主题专家构建模型的速度提高了2.8倍,预测性能比手工标记的7小时平均提高了45.5%。我们研究了这种新环境下的建模权衡,并提出了一种用于自动权衡决策的优化器,该优化器可使每次流水线执行的加速比高达1.8倍。在与美国退伍军人事务部和美国食品和药物管理局的两次合作中,以及在代表其他部署的四个开源文本和图像数据集上,Snorkel比以前的启发式方法提供了132%的平均预测性能改进,并且处于大型手动培训集预测性能的平均3.60%以内。
Labeling training data is increasingly the largest bottleneck in deploying machine learning systems. We present Snorkel, a first-of-its-kind system that enables users to train state-of- the-art models without hand labeling any training data. Instead, users write labeling functions that express arbitrary heuristics, which can have unknown accuracies and correlations. Snorkel denoises their outputs without access to ground truth by incorporating the first end-to-end implementation of our recently proposed machine learning paradigm, data programming. We present a flexible interface layer for writing labeling functions based on our experience over the past year collaborating with companies, agencies, and research labs. In a user study, subject matter experts build models 2.8× faster and increase predictive performance an average 45.5% versus seven hours of hand labeling. We study the modeling tradeoffs in this new setting and propose an optimizer for automating tradeoff decisions that gives up to 1.8× speedup per pipeline execution. In two collaborations, with the U.S. Department of Veterans Affairs and the U.S. Food and Drug Administration, and on four open-source text and image data sets representative of other deployments, Snorkel provides 132% average improvements to predictive performance over prior heuristic approaches and comes within an average 3.60% of the predictive performance of large hand-curated training sets.
DOI: 10.1093/nar/gkv1164
发表时间: 2016-01-04
影响因子: 14.9
作者:
Caspi R;Billington R;Ferrer L;Foerster H;Fulcher CA;Keseler IM;Kothari A;Krummenacker M;Latendresse M;Mueller LA;Ong Q;Paley S;Subhraveti P;Weaver DS;Karp PD
通讯作者: Karp PD
DOI: 10.1109/tit.1970.1054472
发表时间: 1970-01-01
影响因子: 2.5
作者:
AGRAWALA, AK
通讯作者: AGRAWALA, AK
DOI: 10.1093/database/bat080
发表时间: 2013
期刊: Database : the journal of biological databases and curation
影响因子: --
作者:
Davis AP;Wiegers TC;Roberts PM;King BL;Lay JM;Lennon-Hopkins K;Sciaky D;Johnson R;Keating H;Greene N;Hernandez R;McConnell KJ;Enayetallah AE;Mattingly CJ
通讯作者: Mattingly CJ
DOI: 10.1016/j.neunet.2005.06.042
发表时间: 2005-06-01
期刊: NEURAL NETWORKS
影响因子: 7.8
作者:
Graves, A;Schmidhuber, J
通讯作者: Schmidhuber, J