Compressing Word Embeddings via Deep Compositional Code Learning

Compressing Word Embeddings via Deep Compositional Code Learning
复制标题

DOI:
--
复制
发表时间:
2017-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Raphael Shu;Hideki Nakayama
Raphael Shu;Hideki Nakayama
中科院分区:
其他
文献类型:
--
作者:
Raphael Shu;Hideki Nakayama

文献摘要

相似文献

自然语言处理 (NLP) 模型通常需要大量参数进行词嵌入,从而导致大量存储或内存占用。将神经 NLP 模型部署到移动设备需要压缩词嵌入,而不会显着牺牲性能。为此,我们建议用很少的基向量构建嵌入。对于每个单词,基本向量的组成由哈希码确定。为了最大化压缩率,我们采用多码本量化方法而不是二进制编码方案。每个代码由多个离散数字组成,例如(3,2,1,8),其中每个组成部分的值被限制在固定范围内。我们建议通过应用 Gumbel-softmax 技巧来直接学习端到端神经网络中的离散代码。实验表明,在情感分析任务中压缩率达到98%,在机器翻译任务中压缩率达到94%~99%,并且没有性能损失。在这两个任务中,所提出的方法可以通过稍微降低压缩率来提高模型性能。与字符级分割等其他方法相比,所提出的方法是独立于语言的,并且不需要修改网络架构。
Natural language processing (NLP) models often require a massive number of parameters for word embeddings, resulting in a large storage or memory footprint. Deploying neural NLP models to mobile devices requires compressing the word embeddings without any significant sacrifices in performance. For this purpose, we propose to construct the embeddings with few basis vectors. For each word, the composition of basis vectors is determined by a hash code. To maximize the compression rate, we adopt the multi-codebook quantization approach instead of binary coding scheme. Each code is composed of multiple discrete numbers, such as (3, 2, 1, 8), where the value of each component is limited to a fixed range. We propose to directly learn the discrete codes in an end-to-end neural network by applying the Gumbel-softmax trick. Experiments show the compression rate achieves 98% in a sentiment analysis task and 94% ~ 99% in machine translation tasks without performance loss. In both tasks, the proposed method can improve the model performance by slightly lowering the compression rate. Compared to other approaches such as character-level segmentation, the proposed method is language-independent and does not require modifications to the network architecture.