An Empirical Survey of Data Augmentation for Limited Data Learning in NLP

An Empirical Survey of Data Augmentation for Limited Data Learning in NLP
复制标题

DOI:
10.1162/tacl_a_00542
复制
发表时间:
2021-06
影响因子:
10.9
通讯作者:
Jiaao Chen;Derek Tam;Colin Raffel;Mohit Bansal;Diyi Yang
Jiaao Chen;Derek Tam;Colin Raffel;Mohit Bansal;Diyi Yang
中科院分区:
人文科学1区
文献类型:
--
作者:
Jiaao Chen;Derek Tam;Colin Raffel;Mohit Bansal;Diyi Yang

文献摘要

被引文献

相似文献

NLP 在过去十年中通过使用神经模型和大型标记数据集取得了巨大进步。对丰富数据的依赖阻碍了 NLP 模型应用于资源匮乏的环境或需要大量时间、金钱或专业知识来标记大量文本数据的新颖任务。最近,数据增强方法被探索作为提高 NLP 数据效率的一种手段。迄今为止,在有限的标记数据设置中,还没有对 NLP 的数据增强进行系统的实证概述,因此很难理解哪些方法在哪些设置中有效。在本文中,我们对有限标记数据设置下 NLP 数据增强的最新进展进行了实证调查,总结了方法的概况(包括标记级增强、句子级增强、对抗性增强和隐藏空间增强),并在涵盖主题/新闻分类、推理任务、释义任务和单句任务的 11 个数据集上进行了实验。根据结果​​,我们得出几个结论,以帮助从业者在不同的环境中选择适当的增强,并讨论 NLP 中有限数据学习的当前挑战和未来方向。
NLP has achieved great progress in the past decade through the use of neural models and large labeled datasets. The dependence on abundant data prevents NLP models from being applied to low-resource settings or novel tasks where significant time, money, or expertise is required to label massive amounts of textual data. Recently, data augmentation methods have been explored as a means of improving data efficiency in NLP. To date, there has been no systematic empirical overview of data augmentation for NLP in the limited labeled data setting, making it difficult to understand which methods work in which settings. In this paper, we provide an empirical survey of recent progress on data augmentation for NLP in the limited labeled data setting, summarizing the landscape of methods (including token-level augmentations, sentence-level augmentations, adversarial augmentations, and hidden-space augmentations) and carrying out experiments on 11 datasets covering topics/news classification, inference tasks, paraphrasing tasks, and single-sentence tasks. Based on the results, we draw several conclusions to help practitioners choose appropriate augmentations in different settings and discuss the current challenges and future directions for limited data learning in NLP.