Pitfalls of Static Language Modelling

Pitfalls of Static Language Modelling
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
Angeliki Lazaridou;A. Kuncoro;E. Gribovskaya;Devang Agrawal;Adam Liska;Tayfun Terzi;Mai Gimenez;Cyprien de Masson d'Autume;Sebastian Ruder;Dani Yogatama;Kris Cao;Tomás Kociský;Susannah Young;Phil Blunsom
Angeliki Lazaridou;A. Kuncoro;E. Gribovskaya;Devang Agrawal;Adam Liska;Tayfun Terzi;Mai Gimenez;Cyprien de Masson d'Autume;Sebastian Ruder;Dani Yogatama;Kris Cao;Tomás Kociský;Susannah Young;Phil Blunsom
中科院分区:
其他
文献类型:
--
作者:
Angeliki Lazaridou;A. Kuncoro;E. Gribovskaya;Devang Agrawal;Adam Liska;Tayfun Terzi;Mai Gimenez;Cyprien de Masson d'Autume;Sebastian Ruder;Dani Yogatama;Kris Cao;Tomás Kociský;Susannah Young;Phil Blunsom

文献摘要

被引文献

相似文献

我们的世界是开放的,非静止的,不断发展的;因此,我们谈论的东西和我们谈论它的方式随着时间的推移而变化。语言的这种固有的动态性质与当前的静态语言建模范式形成鲜明对比,后者从重叠的时间段构建训练和评估集。尽管最近取得了进展,我们证明了最先进的Transformer模型在预测未来话语的现实设置中表现更差,超出了训练期-来自两个域的三个数据集的一致模式。我们发现,虽然单独增加模型大小-最近进展背后的关键驱动因素-并不能为时间泛化问题提供解决方案,但使用新信息不断更新其知识的模型确实可以随着时间的推移减缓退化。因此,考虑到越来越大的语言建模训练数据集的编译,再加上越来越多的基于语言模型的NLP应用程序需要关于世界的最新知识,我们认为现在是重新思考我们的静态语言建模评估协议的正确时机,并开发自适应语言模型,以便在我们不断变化和非静止的世界中保持最新。1
Our world is open-ended, non-stationary and constantly evolving; thus what we talk about and how we talk about it changes over time. This inherent dynamic nature of language comes in stark contrast to the current static language modelling paradigm, which constructs training and evaluation sets from overlapping time periods. Despite recent progress, we demonstrate that state-of-the-art Transformer models perform worse in the realistic setup of predicting future utterances from beyond their training period—a consistent pattern across three datasets from two domains. We find that, while increasing model size alone—a key driver behind recent progress—does not provide a solution for the temporal generalization problem, having models that continually update their knowledge with new information can indeed slow down the degradation over time. Hence, given the compilation of ever-larger language modelling training datasets, combined with the growing list of language-model-based NLP applications that require up-to-date knowledge about the world, we argue that now is the right time to rethink our static language modelling evaluation protocol, and develop adaptive language models that can remain up-to-date with respect to our ever-changing and non-stationary world. 1