Coarse-grained reconfigurable architectures for machine learning applications
Coarse-grained reconfigurable architectures for machine learning applications
批准号:
1941039
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --
中文摘要
神经网络正在成为解决计算机视觉问题的最新技术,它们在大规模图像分类中取得了出色的准确性。大多数现有的卷积神经网络都很复杂:它们需要大量的参数和密集的计算才能在大型数据集上实现高精度。因此,在嵌入式硬件上运行大型网络有两个主要困难。首先,神经网络的大量参数需要大量的硬件内存。其次,能量消耗主要是内存访问,访问神经网络的大量参数超出了功率敏感嵌入式系统的能量包络。为了解决上述问题,计算机体系结构界目前正在探索神经网络推理的新型硬件体系结构。许多新的定制硬件架构被提出用于神经网络推理和训练。近年来,各种基于fpga的加速器已被应用于神经网络。此外,用于深度神经网络推理和训练的ASIC设计越来越多。这些加速器通常利用一个大的片上存储器,并有定制的计算单元来计算矩阵点积。定制加速器对于有效地运行神经网络计算无疑是有益的,但是访问和存储内存中的大量参数仍然是这些加速器的基本限制。本研究将探索新的网络压缩方法,并开发更有效的数字表示系统,以帮助减少神经网络的大小和计算复杂度。压缩后的神经网络更容易在任何网络加速器上执行。目前,网络压缩领域的研究非常活跃。为了减少神经网络的参数数量,修剪和正则化是常用的方法。网络修剪去除一些不重要的连接或神经元,然后重新训练得到的较小的拓扑结构。如果再训练收敛并且测试精度保持不变,这表明修剪过程成功地找到了更小的网络拓扑。修剪方法可以分为细粒度(修剪单个权重)和粗粒度(修剪整个过滤器)。正则化通过鼓励神经网络在训练阶段的稀缺性来减少神经网络的参数。该研究旨在将剪枝与正则化相结合,以获得更好的压缩效果。此外,研究将侧重于以非确定性的方式修剪,其中先前修剪的权重可以恢复,如果它们后来发现了它们的重要性。减小每个单独参数的位宽度对神经网络的大小有直接影响。常用的数字表示系统,如定点表示和动态定点表示,已经在各种神经网络上得到了充分的发现和评估。这两种算法都是线性量化方法。非线性量化,这种不同的编码方案,在表示权重方面更有效,但由于增加了额外的硬件编码器和解码器,计算复杂性很高。我未来的部分研究将集中在探索新的数字表示系统,以非线性方式利用量化,但保持相对较低的计算成本。
英文摘要
Neural networks are becoming a state-of-the-art technique for solving problems in computer vision, they achieve outstanding accuracies in large-scale image classifications. Most existing convolutional neural networks are complex: they require a large number of parameters and intensive computations to achieve high accuracies on large datasets. Running big networks on embedded hardware therefore has two major difficulties. First, the large number of parameters of a neural network requires a lot of hardware memory. Second, energy consumption is dominated by memory accesses, and accessing the large number of parameters of a neural network exceeds the energy envelope of power sensitive embedded systems. To resolve the above problems, the computer architecture community is currently exploring novel hardware architectures for neural network inference. Many novel custom hardware architectures have been proposed for neural network inference and training. Various FPGA-based accelerators have been recently applied to neural networks. In addition, there is an increasing number of ASIC designs for deep neural network inference and training. These accelerators normally utilize a large on-chip memory and have custom computing units to calculate matrix dot-products. A custom accelerator is definitely beneficial for running neural network computations efficiently, but accessing and storing the large number of parameters in the memory is still a fundamental limit for these accelerators. This research will explore novel network compression methods and exploit more efficient number representation systems to help reduce the size and computation complexity of neural networks. The compressed neural network is then easier to execute on any network accelerators.The area of network compression is now under active research. For reducing the number of parameters of a neural network, pruning and regularization are popular methods. Network pruning removes some unimportant connections or neurons and then retrain the obtained smaller topology. If retraining converges and test accuracy remains unchanged, this suggests that the pruning process successfully finds a smaller network topology. Pruning methods can be classified into fine-grained (pruning individual weights) and coarse-grained (pruning entire filters). Regularization reduces parameters of a neural network by encouraging scarcity in a neural network during the training phase. The research aims to combine pruning with regularization to achieve better compression results. Additionally, the research will focus on pruning in a non-deterministic manner, where previously pruned weights can recover if they found their importance later.Reducing the bit-width of each individual parameters has a direct impact on the size of a neural network. Popular number representation systems such as fixed-point and dynamic fixed-point representations have been fully discovered and evaluated on various neural networks. Both of these proposed arithmetics are linear quantization methods. Non-linear quantizations, such various encoding schemes, are more efficient in representing weights but suffer from high computation complexity by adding extra hardware encoders and decoders. Part of my future research would focus on exploring novel number representation systems that exploit quantization in a non-linear fashion but maintains relatively low computational costs.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.17863/cam.86258
发表时间:
2022
期刊:
影响因子:
--
作者:
[Zhao Y]
通讯作者:
Zhao Y
海外基金