Towards Hybrid Human-AI Workflows for Unknown Unknown Detection

Towards Hybrid Human-AI Workflows for Unknown Unknown Detection
复制标题

迈向用于未知未知检测的混合人类人工智能工作流程

DOI:
--
复制
发表时间:
2020
期刊:
The Web Conference
影响因子:
--
通讯作者:
Walter S. Lasecki
Walter S. Lasecki
中科院分区:
--
文献类型:
--
作者:
Anthony Z. Liu;Santiago Guerra;Isaac Fung;Gabriel Matute;Ece Kamar;Walter S. Lasecki

文献摘要

参考文献

被引文献

相似文献

预测模型容易受到称为未知未知数的错误的影响,其中模型将不正确的标签分配给具有高置信度的实例。当训练数据不代表模型部署时遇到的类的变体时,通常会出现这些问题。先前的工作表明,人群工作者可以识别未知未知的实例,但要求人群识别足够数量的个体实例可能会花费很高的成本[2]。相反,本文提出了一种方法,利用人们发现模式的能力,用更少的例子更有效地重新训练分类器。我们要求人群工作者提出并验证未知的未知模式。然后,我们使用这些模式来训练扩展分类器,以从主分类器过去遇到(并且可能被错误分类)的现有数据中识别其他示例。实验结果表明,该方法在提高分类器性能方面优于现有的未知检测方法。这项工作是第一次利用人群来识别大型数据集中的错误模式,以改进ML训练。
Predictive models are susceptible to errors called unknown unknowns, in which the model assigns incorrect labels to instances with high confidence. These commonly arise when training data does not represent variations of a class encountered at model deployment. Prior work showed that crowd workers can identify instances of unknown unknowns, but asking the crowd to identify a sufficient number of individual instances can be costly to acquire [2]. Instead, this paper presents an approach that leverages people’s ability to find patterns to retrain classifiers more effectively with fewer examples. We ask crowd workers to suggest and verify patterns in unknown unknowns. We then use these patterns to train an expansion classifier to identify additional examples from existing data that the primary classifier has encountered (and potentially misclassified) in the past. Our experiments show that our approach outperforms existing unknown unknown detection methods at improving classifier performance. This work is the first to leverage crowds to identify error patterns in large datasets to improve ML training.
DOI: 10.1609/aaai.v32i1.11493
发表时间: 2018-04
期刊: --
影响因子: --
作者:
Gagan Bansal;Daniel S. Weld
通讯作者: Gagan Bansal;Daniel S. Weld