Optimizing Convolutional Neural Networks on the Sunway TaihuLight Supercomputer

Optimizing Convolutional Neural Networks on the Sunway TaihuLight Supercomputer
复制标题

在神威·太湖之光超级计算机上优化卷积神经网络

DOI:
10.1145/3177885
复制
发表时间:
2018
影响因子:
1.6
通讯作者:
Yang Guangwen
Yang Guangwen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zhao Wenlai;Fu Haohuan;Fang Jiarui;Zheng Weijie;Gan Lin;Yang Guangwen

文献摘要

被引文献

相似文献

神威太湖之光超级计算机由SW26010驱动,SW26010是一款新型260核处理器,采用片上异构核融合设计。在本文中,我们介绍了我们在神威太湖之光超级计算机上优化卷积神经网络(cnn)训练过程的工作。具体来说,提出了一个高效的库(swDNN)和一个定制的Caffe框架(swCaffe)。介绍了针对SW26010多核架构的面向体系结构的优化方法,与未优化的算法和框架相比,使用swCaffe对swDNN中的卷积例程实现了48倍的加速,对VGG-16网络的整个训练过程实现了4倍的加速。与基于NVIDIA K40m GPU的cuDNN库和Caffe框架相比,基于SW26010的swDNN库和Caffe框架在单精度下的性能几乎是K40m的一半,在双精度下的速度分别是K40m的3.6倍和1.8倍。
The Sunway TaihuLight supercomputer is powered by SW26010, a new 260-core processor designed with on-chip fusion of heterogeneous cores. In this article, we present our work on optimizing the training process of convolutional neural networks (CNNs) on the Sunway TaihuLight supercomputer. Specifically, a highly efficient library (swDNN) and a customized Caffe framework (swCaffe) are proposed. Architecture-oriented optimization methods targeting the many-core architecture of SW26010 are introduced and are able to achieve 48× speedup for the convolution routine in swDNN and 4× speedup for the complete training process of the VGG-16 network using swCaffe, compared to the unoptimized algorithm and framework. Compared to the cuDNN library and the Caffe framework based on the NVIDIA K40m GPU, the proposed swDNN library and swCaffe framework on SW26010 have nearly half the performance of K40m in single -precision and have 3.6× and 1.8× speedup over K40m in double precision, respectively.