FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters

FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters
复制标题

DOI:
10.1109/tc.2020.3000118
复制
发表时间:
2020-05
影响因子:
3.7
通讯作者:
Tianqi Wang;Tong Geng;Ang Li;Xi Jin;Martin C. Herbordt
Tianqi Wang;Tong Geng;Ang Li;Xi Jin;Martin C. Herbordt
中科院分区:
计算机科学2区
文献类型:
--
作者:
Tianqi Wang;Tong Geng;Ang Li;Xi Jin;Martin C. Herbordt

文献摘要

被引文献

相似文献

深度卷积神经网络(CNN)已经彻底改变了许多应用,但对更高性能的需求仍然有增无减。将CNN计算扩展到更大的集群通常是通过使用分布式同步SGD等方法以批处理模式分发任务来完成的。这种方法的问题之一是,为了使分布式集群以高利用率工作,分配给每个节点的工作负载必须很大;这意味着SGD小批量大小的增长非常重要。在本文中,我们提出了一个名为FPDeep的框架,它使用模型和层并行的混合来配置分布式可重构集群来训练CNN。这种方法有许多好处。首先,该设计不会由于批量大小的增长而遭受性能损失。其次,工作和存储之间的平衡节点通过新的工作负载和重量划分方案。该机制的一部分是令人惊讶的发现,最好将多余的权重存储在相邻设备中,而不是本地片外存储器中。第三,整个系统是一个细粒度的管道。这导致了高并行性和利用率,并且还最小化了在等待反向传播时需要缓存的功能的时间。因此,存储需求降低到仅将片上存储器用于卷积层的程度。第四,我们发现,最简单的拓扑结构,一个1D阵列,是首选的FPGA互连,从而使广泛的适用性。我们使用Alexnet、VGG-16和VGG-19基准测试来评估FPDeep。结果表明,FPDeep对大量FPGA具有良好的可扩展性,限制因素是FPGA到FPGA的带宽。但是,由于每个FPGA的双向带宽为250 Gb/s,这很容易被当前一代FPGA支持,因此FPDeep的性能表现出高达100个FPGA的线性。能源效率是相对于GOP/J进行评估的。FPDeep提供的能源效率平均比同类GPU服务器高6.4倍。
Deep convolutional Neural Networks (CNNs) have revolutionized numerous applications, but the demand for ever more performance remains unabated. Scaling CNN computations to larger clusters is generally done by distributing tasks in batch mode using methods such as distributed synchronous SGD. Among the issues with this approach is that, to make the distributed cluster work with high utilization, the workload distributed to each node must be large; this implies nontrivial growth in the SGD mini-batch size. In this article we propose a framework, called FPDeep, which uses a hybrid of model and layer parallelism to configure distributed reconfigurable clusters to train CNNs. This approach has numerous benefits. First, the design does not suffer from performance loss due to batch size growth. Second, work and storage are balanced among nodes through novel workload and weight partitioning schemes. Part of the mechanism is the surprising finding that it is preferable to store excess weights in neighboring devices rather than in local off-chip memory. Third, the entire system is a fine-grained pipeline. This leads to high parallelism and utilization and also minimizes the time that features need to be cached while waiting for back-propagation. As a result, storage demand is reduced to the point where only on-chip memory is used for the convolution layers. And fourth, we find that the simplest topology, a 1D array, is preferred for interconnecting the FPGAs thus enabling widespread applicability. We evaluate FPDeep with the Alexnet, VGG-16, and VGG-19 benchmarks. Results show that FPDeep has good scalability to a large number of FPGAs, with the limiting factor being the FPGA-to-FPGA bandwidth. But with 250 Gb/s bidirectional bandwidth per FPGA, which is easily supported by current generation FPGAs, FPDeep performance shows linearity up to 100 FPGAs. Energy efficiency is evaluated with respect to GOPs/J. FPDeep provides, on average, 6.4× higher energy efficiency than comparable GPU servers.