Efficient Network Construction Through Structural Plasticity

Efficient Network Construction Through Structural Plasticity
复制标题

DOI:
10.1109/jetcas.2019.2933233
复制
发表时间:
2019-05
影响因子:
4.6
通讯作者:
Xiaocong Du;Zheng Li;Yufei Ma;Yu Cao
Xiaocong Du;Zheng Li;Yufei Ma;Yu Cao
中科院分区:
工程技术2区
文献类型:
--
作者:
Xiaocong Du;Zheng Li;Yufei Ma;Yu Cao

文献摘要

相似文献

基于硬件的深度神经网络(Deep Neural Networks, dnn)由于参数数量庞大而面临计算成本过高的问题。缓解过度参数化的典型训练管道是,以高精度为目标,预先定义具有冗余学习单元(过滤器和神经元)的DNN结构,然后在训练后修剪冗余学习单元,以实现高效推理。我们认为在训练中引入冗余是次优的,以便在以后的推理中减少冗余。此外,固定的网络结构进一步导致对终身学习等动态任务的适应能力较差。相比之下,结构可塑性在哺乳动物大脑中发挥着不可或缺的作用,以实现紧凑和准确的学习。在整个生命周期中,活跃的连接不断创建,而那些不再重要的连接则退化。在这种观察的启发下,我们提出了一种训练方案,即连续增长和修剪(CGaP),我们从一个小的网络种子开始训练,然后通过添加重要的学习单元来执行连续增长,最后修剪次要的学习单元以进行有效的推理。CGaP生成的推理模型在结构上是稀疏的,在硬件平台上部署时大大降低了推理能力和延迟。在具有代表性的数据集上使用流行的深度神经网络结构,通过在现场可编程门阵列(FPGA)上进行算法仿真和体系结构建模,对CGaP的有效性进行了基准测试。例如,CGaP使CIFAR-10上的ResNet-110的FLOPs、模型尺寸、DRAM访问能量和推理延迟分别降低了63.3%、64.0%、11.8%和40.2%。
Deep Neural Networks (DNNs) on hardware is facing excessive computation cost due to the massive number of parameters. A typical training pipeline to mitigate over-parameterization is to pre-define a DNN structure with redundant learning units (filters and neurons) with the goal of high accuracy, then to prune redundant learning units after training with the purpose of efficient inference. We argue that it is sub-optimal to introduce redundancy into training in order to reduce redundancy later in inference. Moreover, the fixed network structure further results in poor adaption to dynamic tasks, such as lifelong learning. In contrast, structural plasticity plays an indispensable role in mammalian brains to achieve compact and accurate learning. Throughout the lifetime, active connections are continuously created while those that are no longer important are degenerated. Inspired by such observation, we propose a training scheme, namely Continuous Growth and Pruning (CGaP), where we start the training from a small network seed, then literally execute continuous growth by adding important learning units and finally prune secondary ones for efficient inference. The inference model generated from CGaP is sparse in the structure, largely decreasing the inference power and latency when deployed on hardware platforms. With popular DNN structures on representative datasets, the efficacy of CGaP is benchmarked by both algorithmic simulation and architectural modeling on Field-programmable Gate Arrays (FPGA). For example, CGaP decreases the FLOPs, model size, DRAM access energy and inference latency by 63.3%, 64.0%, 11.8% and 40.2%, respectively, for ResNet-110 on CIFAR-10.