Addressing Semantic Drift in Question Generation for Semi-Supervised Question Answering

Addressing Semantic Drift in Question Generation for Semi-Supervised Question Answering
复制标题

DOI:
10.18653/v1/d19-1253
复制
发表时间:
2019-09
期刊:
--
影响因子:
--
通讯作者:
Shiyue Zhang;Mohit Bansal
Shiyue Zhang;Mohit Bansal
中科院分区:
其他
文献类型:
--
作者:
Shiyue Zhang;Mohit Bansal

文献摘要

被引文献

相似文献

基于文本的问题生成(Text-Based Query Generation,QG)旨在生成自然的、相关的问题,这些问题可以在一定的上下文中由给定的答案来回答。现有的QG模型存在“语义漂移”问题,即模型生成的问题的语义偏离了给定的上下文和答案。在本文中,我们首先提出了两个语义增强的奖励,从下游问题释义和问答任务中获得,以规范QG模型,生成语义有效的问题。其次,针对传统的评价指标(如BLEU)在评价问题生成质量方面的不足,我们提出了一种基于QA的评价方法,该方法衡量了QG模型在生成QA训练数据时模仿人类注释者的能力。实验表明,我们的方法达到了最新的性能W.r.t.传统指标,在我们基于质量保证的评估指标上也表现最好。此外,我们还研究了如何使用我们的QG模型来扩充QA数据集和实现半监督QA。我们提出了两种生成合成QA对的方法:从现有文章生成新问题或从新文章收集QA对。我们还提出了两种经验上有效的策略,数据过滤器和混合小批量训练,以适当地使用QG生成的数据进行QA。实验表明,即使在不引入新文章的情况下,我们的方法也比BiDAF和BERT QA基线都有所改善。
Text-based Question Generation (QG) aims at generating natural and relevant questions that can be answered by a given answer in some context. Existing QG models suffer from a “semantic drift” problem, i.e., the semantics of the model-generated question drifts away from the given context and answer. In this paper, we first propose two semantics-enhanced rewards obtained from downstream question paraphrasing and question answering tasks to regularize the QG model to generate semantically valid questions. Second, since the traditional evaluation metrics (e.g., BLEU) often fall short in evaluating the quality of generated questions, we propose a QA-based evaluation method which measures the QG model’s ability to mimic human annotators in generating QA training data. Experiments show that our method achieves the new state-of-the-art performance w.r.t. traditional metrics, and also performs best on our QA-based evaluation metrics. Further, we investigate how to use our QG model to augment QA datasets and enable semi-supervised QA. We propose two ways to generate synthetic QA pairs: generate new questions from existing articles or collect QA pairs from new articles. We also propose two empirically effective strategies, a data filter and mixing mini-batch training, to properly use the QG-generated data for QA. Experiments show that our method improves over both BiDAF and BERT QA baselines, even without introducing new articles.