Calibration, Entropy Rates, and Memory in Language Models

Calibration, Entropy Rates, and Memory in Language Models
复制标题

DOI:
--
复制
发表时间:
2019-06
期刊:
Proceedings of the conference. Association for Computational Linguistics. North American Chapter. Meeting
影响因子:
--
通讯作者:
M. Braverman;Xinyi Chen;S. Kakade;Karthik Narasimhan;Cyril Zhang;Yi Zhang
M. Braverman;Xinyi Chen;S. Kakade;Karthik Narasimhan;Cyril Zhang;Yi Zhang
中科院分区:
其他
文献类型:
--
作者:
M. Braverman;Xinyi Chen;S. Kakade;Karthik Narasimhan;Cyril Zhang;Yi Zhang

文献摘要

相似文献

构建能够捕捉有意义的长期依赖关系的准确语言模型是自然语言处理中的一个核心挑战。为此,我们提出了一种基于校准的方法来衡量生成式序列模型与真实分布之间的长期差异,并利用这些差异来改进模型。从经验上看,我们表明,包括长短期记忆网络(LSTMs)和变换器(Transformers)在内的最先进的语言模型是\emph{校准不当的}:它们生成内容的熵率随着时间推移急剧上升。然后,我们提供了可证明的方法来缓解这种现象。此外,我们还展示了这种基于校准的方法如何也可用于衡量语言模型用于预测的记忆量。
Building accurate language models that capture meaningful long-term dependencies is a core challenge in natural language processing. Towards this end, we present a calibration-based approach to measure long-term discrepancies between a generative sequence model and the true distribution, and use these discrepancies to improve the model. Empirically, we show that state-of-the-art language models, including LSTMs and Transformers, are \emph{miscalibrated}: the entropy rates of their generations drift dramatically upward over time. We then provide provable methods to mitigate this phenomenon. Furthermore, we show how this calibration-based approach can also be used to measure the amount of memory that language models use for prediction.