Leveraging QA Datasets to Improve Generative Data Augmentation

Leveraging QA Datasets to Improve Generative Data Augmentation
复制标题

DOI:
10.18653/v1/2022.emnlp-main.660
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Dheeraj Mekala;Tu Vu;Timo Schick;Jingbo Shang
Dheeraj Mekala;Tu Vu;Timo Schick;Jingbo Shang
中科院分区:
其他
文献类型:
--
作者:
Dheeraj Mekala;Tu Vu;Timo Schick;Jingbo Shang

文献摘要

被引文献

相似文献

生成语言模型(glm)生成文本的能力在过去几年中有了很大的改进,使它们能够用于生成数据增强。在这项工作中,我们提出了CONDA,这是一种进一步提高GLM生成合成数据的能力的方法,通过将数据生成重新表述为给定问答(QA)对的上下文生成,并利用QA数据集来训练上下文生成器。然后,我们将下游任务转换为相同的问答格式,并使经过微调的上下文生成器适应目标任务域。最后,我们使用微调的GLM来生成相关上下文,这些上下文反过来用作相应任务的综合训练数据。我们在多个分类数据集上进行了广泛的实验,并证明了在少量和零射击设置下性能的实质性改进。我们的分析表明,需要高级推理能力的QA数据集(例如,抽象和常识性QA数据集)倾向于在少量射击和零射击设置中提供最好的性能提升。
The ability of generative language models (GLMs) to generate text has improved considerably in the last few years, enabling their use for generative data augmentation. In this work, we propose CONDA, an approach to further improve GLM’s ability to generate synthetic data by reformulating data generation as context generation for a given question-answer (QA) pair and leveraging QA datasets for training context generators. Then, we cast downstream tasks into the same question answering format and adapt the fine-tuned context generators to the target task domain. Finally, we use the fine-tuned GLM to generate relevant contexts, which are in turn used as synthetic training data for their corresponding tasks. We perform extensive experiments on multiple classification datasets and demonstrate substantial improvements in performance for both few- and zero-shot settings. Our analysis reveals that QA datasets that require high-level reasoning abilities (e.g., abstractive and common-sense QA datasets) tend to give the best boost in performance in both few-shot and zero-shot settings.