Life after BERT: What do Other Muppets Understand about Language?

Life after BERT: What do Other Muppets Understand about Language?
复制标题

DOI:
10.18653/v1/2022.acl-long.227
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Vladislav Lialin;Kevin Zhao;Namrata Shivagunde;Anna Rumshisky
Vladislav Lialin;Kevin Zhao;Namrata Shivagunde;Anna Rumshisky
中科院分区:
其他
文献类型:
--
作者:
Vladislav Lialin;Kevin Zhao;Namrata Shivagunde;Anna Rumshisky

文献摘要

相似文献

现有的预训练的Transformer分析工作通常一次只关注一个或两个模型族,忽略了架构和预训练目标的可变性。在我们的工作中,我们利用oLMpics基准和心理语言学探测数据集,用于包括T5,BART和ALBERT在内的29个模型。此外,我们还将oLMpics零触发设置用于自回归模型,并评估不同大小的GPT网络。我们的研究结果表明,这些模型中没有一个可以以零射击的方式解决组合问题,这表明这种技能是无法使用现有的预训练目标来学习的。此外,我们发现,全局模型决策,如架构,方向性,数据集大小和预训练目标并不能预测模型的语言能力。
Existing pre-trained transformer analysis works usually focus only on one or two model families at a time, overlooking the variability of the architecture and pre-training objectives. In our work, we utilize the oLMpics bench- mark and psycholinguistic probing datasets for a diverse set of 29 models including T5, BART, and ALBERT. Additionally, we adapt the oLMpics zero-shot setup for autoregres- sive models and evaluate GPT networks of different sizes. Our findings show that none of these models can resolve compositional questions in a zero-shot fashion, suggesting that this skill is not learnable using existing pre-training objectives. Furthermore, we find that global model decisions such as architecture, directionality, size of the dataset, and pre-training objective are not predictive of a model’s linguistic capabilities.