Direct Output Connection for a High-Rank Language Model

Direct Output Connection for a High-Rank Language Model
复制标题

DOI:
10.18653/v1/d18-1489
复制
发表时间:
2018-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Sho Takase;Jun Suzuki;M. Nagata
Sho Takase;Jun Suzuki;M. Nagata
中科院分区:
其他
文献类型:
--
作者:
Sho Takase;Jun Suzuki;M. Nagata

文献摘要

被引文献

相似文献

本文提出了一种最先进的递归神经网络(RNN)语言模型,它结合了不仅从最终RNN层而且从中间层计算的概率分布。这种方法基于Yang等人(2018)引入的语言建模的矩阵分解解释,提高了语言模型的表达能力。我们提出的方法改进了当前最先进的语言模型,并在标准基准数据集Penn Treebank和WikiText-2上获得了最佳分数。此外,我们指出,我们提出的方法有助于应用任务:机器翻译和标题生成。
This paper proposes a state-of-the-art recurrent neural network (RNN) language model that combines probability distributions computed not only from a final RNN layer but also middle layers. This method raises the expressive power of a language model based on the matrix factorization interpretation of language modeling introduced by Yang et al. (2018). Our proposed method improves the current state-of-the-art language model and achieves the best score on the Penn Treebank and WikiText-2, which are the standard benchmark datasets. Moreover, we indicate our proposed method contributes to application tasks: machine translation and headline generation.