Non-uniform DNN Structured Subnets Sampling for Dynamic Inference

Non-uniform DNN Structured Subnets Sampling for Dynamic Inference
复制标题

DOI:
10.1109/dac18072.2020.9218736
复制
发表时间:
2020-07
期刊:
2020 57th ACM/IEEE Design Automation Conference (DAC)
影响因子:
--
通讯作者:
Li Yang-;Zhezhi He;Yu Cao;Deliang Fan
Li Yang-;Zhezhi He;Yu Cao;Deliang Fan
中科院分区:
其他
文献类型:
--
作者:
Li Yang-;Zhezhi He;Yu Cao;Deliang Fan

文献摘要

相似文献

随着深度神经网络(DNN)的成功,许多最近的作品一直集中在通过模型压缩技术(例如量化,修剪,近距离近似等等)来开发电源和资源有限系统的硬件加速器,等等。但是,几乎所有现有的压缩DNN在部署后是固定的,该DNN缺乏运行时自适应结构,无法适应其动态硬件资源分配,功率预算,吞吐量需求以及动态工作负载。作为对策,为了构建一种新型的运行时动态DNN结构,我们通过子网生成的非均匀通道选择提出了一种新型的DNN子网络采样方法。因此,根据给定系统的动态要求或规格,用户可以在部署后的功率,速度,计算负载和准确性之间进行权衡。我们使用Resnets验证了CIFAR-10和Imagenet数据集上提出的模型,该模型的表现优于接受单独训练的子网和其他相关作品。它表明,我们的方法可以在GPU的13.4、24.6、41.3、62.1(MS)和30.5、38.7、51、65.4(MS)的13.4、24.6、41.3、62.1(MS)和38.5、38.7、51、65.4(MS)中,分别使用Resnet18在ImagEnet上分别在Imagenet上进行了128批批次和CPU。
With the success of Deep Neural Networks (DNN), many recent works have been focusing on developing hardware accelerator for power and resource-limited system via model compression techniques, such as quantization, pruning, low-rank approximation and etc. However, almost all existing compressed DNNs are fixed after deployment, which lacks run-time adaptive structure to adapt to its dynamic hardware resource allocation, power budget, throughput requirement, as well as dynamic workload. As the countermeasure, to construct a novel run-time dynamic DNN structure, we propose a novel DNN sub-network sampling method via non-uniform channel selection for subnets generation. Thus, user can trade off between power, speed, computing load and accuracy on-the-fly after the deployment, depending on the dynamic requirements or specifications of the given system. We verify the proposed model on both CIFAR-10 and ImageNet dataset using ResNets, which outperforms the same sub-nets trained individually and other related works. It shows that, our method can achieve latency trade-off among 13.4, 24.6, 41.3, 62.1(ms) and 30.5, 38.7, 51, 65.4(ms) for GPU with 128 batch-size and CPU respectively on ImageNet using ResNet18.