Improving Robustness by Augmenting Training Sentences with Predicate-Argument Structures

Improving Robustness by Augmenting Training Sentences with Predicate-Argument Structures
复制标题

通过使用谓词-论元结构增强训练句子来提高鲁棒性

DOI:
--
复制
发表时间:
2020
期刊:
arXiv.org
影响因子:
--
通讯作者:
Iryna Gurevych
Iryna Gurevych
中科院分区:
--
文献类型:
--
作者:
N. Moosavi;M. Boer;Prasetya Ajie Utama;Iryna Gurevych

文献摘要

参考文献

被引文献

相似文献

现有的NLP数据集包含各种偏差,并且模型倾向于快速学习这些偏差,这反过来限制了它们的鲁棒性。现有的提高对数据集偏差鲁棒性的方法主要集中在改变训练目标上,这样模型就可以从有偏差的例子中学到更少的东西。此外,它们大多专注于解决特定的偏差,虽然它们提高了目标偏差的对抗性评估集的性能,但它们可能会以其他方式使模型产生偏差,从而损害整体的鲁棒性。在本文中,我们提出用相应的谓词-参数结构来增强训练数据中的输入句子,这为相同含义的不同实现提供了更高层次的抽象,并帮助模型识别句子的重要部分。我们表明,在不针对特定偏差的情况下,我们的句子增强提高了变压器模型对多个偏差的鲁棒性。此外,我们表明,即使在训练数据不包含词汇重叠偏差的情况下,模型仍然容易受到这种偏差的影响,并且句子增强也提高了这种情况下的鲁棒性。我们将在此https URL上发布我们的对抗性数据集来评估这种场景中的偏见以及我们的增强脚本。
Existing NLP datasets contain various biases, and models tend to quickly learn those biases, which in turn limits their robustness. Existing approaches to improve robustness against dataset biases mostly focus on changing the training objective so that models learn less from biased examples. Besides, they mostly focus on addressing a specific bias, and while they improve the performance on adversarial evaluation sets of the targeted bias, they may bias the model in other ways, and therefore, hurt the overall robustness. In this paper, we propose to augment the input sentences in the training data with their corresponding predicate-argument structures, which provide a higher-level abstraction over different realizations of the same meaning and help the model to recognize important parts of sentences. We show that without targeting a specific bias, our sentence augmentation improves the robustness of transformer models against multiple biases. In addition, we show that models can still be vulnerable to the lexical overlap bias, even when the training data does not contain this bias, and that the sentence augmentation also improves the robustness in this scenario. We will release our adversarial datasets to evaluate bias in such a scenario as well as our augmentation scripts at this https URL.
DOI: 10.18653/v1/2020.acl-main.770
发表时间: 2020-05
期刊: --
影响因子: --
作者:
Prasetya Ajie Utama;N. Moosavi;Iryna Gurevych
通讯作者: Prasetya Ajie Utama;N. Moosavi;Iryna Gurevych
DOI: 10.18653/v1/2020.emnlp-main.613
发表时间: 2020-09
期刊: --
影响因子: --
作者:
Prasetya Ajie Utama;N. Moosavi;Iryna Gurevych
通讯作者: Prasetya Ajie Utama;N. Moosavi;Iryna Gurevych