Are Larger Pretrained Language Models Uniformly Better? Comparing Performance at the Instance Level

Are Larger Pretrained Language Models Uniformly Better? Comparing Performance at the Instance Level
复制标题

DOI:
10.18653/v1/2021.findings-acl.334
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Ruiqi Zhong;Dhruba Ghosh;D. Klein;J. Steinhardt
Ruiqi Zhong;Dhruba Ghosh;D. Klein;J. Steinhardt
中科院分区:
其他
文献类型:
--
作者:
Ruiqi Zhong;Dhruba Ghosh;D. Klein;J. Steinhardt

文献摘要

相似文献

较大的语言模型平均具有更高的准确性,但它们在每个实例(数据点)上都更好吗?一些工作表明,较大的模型具有较高的非分布稳健性,而另一些工作表明,它们对稀有子组的精度较低。为了理解这些差异,我们在单个实例的级别上调查这些模型。然而,一个主要的挑战是个体预测在训练的随机性中对噪声高度敏感。我们开发了严格的统计方法来解决这个问题,在考虑了预训练和精调噪声后,我们发现在MNLI、SST-2和QQP上,我们的BERT-Large至少在1%-4%的实例上比BERT-Mini差,而总体准确率提高了2%-10%。我们还发现,精调噪声随着模型大小的增加而增加,并且实例级精度具有动量:从BERT-Mini到BERT-Medium的改进与从BERT-Medium到BERT-Large的改进相关。我们的发现表明,实例级预测提供了丰富的信息来源;因此,我们建议研究人员用模型预测补充模型权重。
Larger language models have higher accuracy on average, but are they better on every single instance (datapoint)? Some work suggests larger models have higher out-of-distribution robustness, while other work suggests they have lower accuracy on rare subgroups. To understand these differences, we investigate these models at the level of individual instances. However, one major challenge is that individual predictions are highly sensitive to noise in the randomness in training. We develop statistically rigorous methods to address this, and after accounting for pretraining and finetuning noise, we find that our BERT-Large is worse than BERT-Mini on at least 1-4% of instances across MNLI, SST-2, and QQP, compared to the overall accuracy improvement of 2-10%. We also find that finetuning noise increases with model size and that instance-level accuracy has momentum: improvement from BERT-Mini to BERT-Medium correlates with improvement from BERT-Medium to BERT-Large. Our findings suggest that instance-level predictions provide a rich source of information; we therefore, recommend that researchers supplement model weights with model predictions.