Neural Networks Weights Quantization: Target None-retraining Ternary (TNT)

Neural Networks Weights Quantization: Target None-retraining Ternary (TNT)
复制标题

DOI:
10.1109/emc2-nips53020.2019.00022
复制
发表时间:
2019-12
期刊:
2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing - NeurIPS Edition (EMC2-NIPS)
影响因子:
--
通讯作者:
Tianyu Zhang;Lei Zhu;Qian Zhao;Kilho Shin
Tianyu Zhang;Lei Zhu;Qian Zhao;Kilho Shin
中科院分区:
其他
文献类型:
--
作者:
Tianyu Zhang;Lei Zhu;Qian Zhao;Kilho Shin

文献摘要

相似文献

深度神经网络(DNN)的权重量化已被证明是在移动设备、asic和fpga等边缘设备上实现DNN的有效解决方案,因为它们没有足够的资源来支持涉及数百万高精度权重和乘法累加运算的计算。本文提出了一种将dnn的高精度权值向量压缩为三元向量的新方法,即基于余弦相似度的目标非再训练三元(TNT)压缩方法。我们的方法利用余弦相似度而不是文献中常用的欧几里得距离,并成功地减少了搜索空间的大小,以找到从$3^{N}$到$N$的最优三元向量,其中$N$是目标向量的维数。因此,TNT寻找理论上最优三元向量的计算复杂度仅为$O(N\log(N))$。此外,我们的实验表明,当我们对具有高精度参数的深度神经网络模型进行ter化时,得到的量化模型可以具有足够高的精度,从而无需重新训练模型。
Quantization of weights of deep neural networks (DNN) has proven to be an effective solution for the purpose of implementing DNNs on edge devices such as mobiles, ASICs and FPGAs, because they have no sufficient resources to support computation involving millions of high precision weights and multiply-accumulate operations. This paper proposes a novel method to compress vectors of high precision weights of DNNs to ternary vectors, namely a cosine similarity based target non-retraining ternary (TNT) compression method. Our method leverages cosine similarity instead of Euclidean distances as commonly used in the literature and succeeds in reducing the size of the search space to find optimal ternary vectors from $3^{N}$ to $N$, where $N$ is the dimension of target vectors. As a result, the computational complexity for TNT to find theoretically optimal ternary vectors is only $O(N\log(N))$. Moreover, our experiments show that, when we ternarize models of DNN with high precision parameters, the obtained quantized models can exhibit sufficiently high accuracy so that re-training models is not necessary.