Optimizing Convolutional Neural Networks on the Sunway TaihuLight Supercomputer
Optimizing Convolutional Neural Networks on the Sunway TaihuLight Supercomputer
复制标题
在神威·太湖之光超级计算机上优化卷积神经网络
DOI:
10.1145/3177885
复制
发表时间:
2018
影响因子:
1.6
通讯作者:
Yang Guangwen
中科院分区:
文献类型:
--
作者:
Zhao Wenlai;Fu Haohuan;Fang Jiarui;Zheng Weijie;Gan Lin;Yang Guangwen
The Sunway TaihuLight supercomputer is powered by SW26010, a new 260-core processor designed with on-chip fusion of heterogeneous cores. In this article, we present our work on optimizing the training process of convolutional neural networks (CNNs) on the Sunway TaihuLight supercomputer. Specifically, a highly efficient library (swDNN) and a customized Caffe framework (swCaffe) are proposed. Architecture-oriented optimization methods targeting the many-core architecture of SW26010 are introduced and are able to achieve 48× speedup for the convolution routine in swDNN and 4× speedup for the complete training process of the VGG-16 network using swCaffe, compared to the unoptimized algorithm and framework. Compared to the cuDNN library and the Caffe framework based on the NVIDIA K40m GPU, the proposed swDNN library and swCaffe framework on SW26010 have nearly half the performance of K40m in single -precision and have 3.6× and 1.8× speedup over K40m in double precision, respectively.