On the Importance of Adaptive Data Collection for Extremely Imbalanced Pairwise Tasks

On the Importance of Adaptive Data Collection for Extremely Imbalanced Pairwise Tasks
复制标题

论自适应数据收集对于极其不平衡的成对任务的重要性

DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Percy Liang
Percy Liang
中科院分区:
--
文献类型:
--
作者:
Stephen Mussmann;Robin Jia;Percy Liang

文献摘要

被引文献

相似文献

许多成对分类任务,例如释义检测和开放域问题回答,自然具有极端的标签不平衡(例如,99.99%的例子是否定的)。相比之下,许多最近的数据集选择示例以确保标签平衡。我们发现,这些算法导致训练的模型泛化能力差:在QQP和WikiQA上训练的最先进的模型在现实不平衡的测试数据上进行评估时,平均精度只有2.4%。相反,我们通过主动学习收集训练数据,使用基于BERT的嵌入模型从大量未标记的话语对中有效地检索不确定点。通过创建具有更多信息的负面示例的平衡训练数据,主动学习大大提高了QQP的平均精度,达到32.5%,WikiQA为20.1%。
Many pairwise classification tasks, such as paraphrase detection and open-domain question answering, naturally have extreme label imbalance (e.g., 99.99% of examples are negatives). In contrast, many recent datasets heuristically choose examples to ensure label balance. We show that these heuristics lead to trained models that generalize poorly: State-of-the art models trained on QQP and WikiQA each have only 2.4% average precision when evaluated on realistically imbalanced test data. We instead collect training data with active learning, using a BERT-based embedding model to efficiently retrieve uncertain points from a very large pool of unlabeled utterance pairs. By creating balanced training data with more informative negative examples, active learning greatly improves average precision to 32.5% on QQP and 20.1% on WikiQA.