When Choosing Plausible Alternatives, Clever Hans can be Clever

When Choosing Plausible Alternatives, Clever Hans can be Clever
复制标题

DOI:
10.18653/v1/d19-6004
复制
发表时间:
2019-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Pride Kavumba;Naoya Inoue;Benjamin Heinzerling;Keshav Singh;Paul Reisert;Kentaro Inui
Pride Kavumba;Naoya Inoue;Benjamin Heinzerling;Keshav Singh;Paul Reisert;Kentaro Inui
中科院分区:
其他
文献类型:
--
作者:
Pride Kavumba;Naoya Inoue;Benjamin Heinzerling;Keshav Singh;Paul Reisert;Kentaro Inui

文献摘要

相似文献

伯特(Bert)和罗伯塔(Roberta)等审计的语言模型在平价推理基准杯中表现出很大的改进,但是,最近的工作发现,自然语言理解的基准的许多改进并不是由于模型学习了这项任务,而是由于他们的能力提高要利用浅表提示,例如在正确的答案中,贝特和罗伯塔在COPA上的良好表现也是如此,我们在COPA中也造成了肤浅的提示。为了解决这个问题,我们介绍了平衡的杯子,这是一个不受易于探索的单一代币提示的阶段,我们分析了Bert和Roberta在原始和平衡的COPA上的表现,发现Bert依靠浅表线索存在,但一旦使它们变得无效,它们仍然可以达到可比的性能,这表明伯特在被迫时会在一定程度上学习任务。
Pretrained language models, such as BERT and RoBERTa, have shown large improvements in the commonsense reasoning benchmark COPA. However, recent work found that many improvements in benchmarks of natural language understanding are not due to models learning the task, but due to their increasing ability to exploit superficial cues, such as tokens that occur more often in the correct answer than the wrong one. Are BERT’s and RoBERTa’s good performance on COPA also caused by this? We find superficial cues in COPA, as well as evidence that BERT exploits these cues.To remedy this problem, we introduce Balanced COPA, an extension of COPA that does not suffer from easy-to-exploit single token cues. We analyze BERT’s and RoBERTa’s performance on original and Balanced COPA, finding that BERT relies on superficial cues when they are present, but still achieves comparable performance once they are made ineffective, suggesting that BERT learns the task to a certain degree when forced to. In contrast, RoBERTa does not appear to rely on superficial cues.