Multiscale sequence modeling with a learned dictionary

Multiscale sequence modeling with a learned dictionary
复制标题

使用学习字典进行多尺度序列建模

DOI:
--
复制
发表时间:
2017
期刊:
arXiv.org
影响因子:
--
通讯作者:
Yoshua Bengio
Yoshua Bengio
中科院分区:
--
文献类型:
--
作者:
B. V. Merrienboer;Amartya Sanyal;H. Larochelle;Yoshua Bengio

文献摘要

被引文献

相似文献

我们提出了一种神经网络序列模型的推广。我们的多尺度模型不是一次预测一个符号,而是对多个可能重叠的多符号标记进行预测。使用字节对编码(BPE)压缩算法的一种变体来学习模型所使用的标记字典。当应用于语言建模时,我们的模型具有字符级模型的灵活性,同时保持了词级模型的许多性能优势。我们的实验表明,该模型在语言建模任务上比常规的长短期记忆网络(LSTM)表现更好,尤其是对于较小的模型。
We propose a generalization of neural network sequence models. Instead of predicting one symbol at a time, our multi-scale model makes predictions over multiple, potentially overlapping multi-symbol tokens. A variation of the byte-pair encoding (BPE) compression algorithm is used to learn the dictionary of tokens that the model is trained with. When applied to language modelling, our model has the flexibility of character-level models while maintaining many of the performance benefits of word-level models. Our experiments show that this model performs better than a regular LSTM on language modeling tasks, especially for smaller models.