Supporting Very Large Models using Automatic Dataflow Graph Partitioning

Supporting Very Large Models using Automatic Dataflow Graph Partitioning
复制标题

DOI:
10.1145/3302424.3303953
复制
发表时间:
2018-07
期刊:
Proceedings of the Fourteenth EuroSys Conference 2019
影响因子:
--
通讯作者:
Minjie Wang;Chien-chin Huang;Jinyang Li
Minjie Wang;Chien-chin Huang;Jinyang Li
中科院分区:
其他
文献类型:
--
作者:
Minjie Wang;Chien-chin Huang;Jinyang Li

文献摘要

被引文献

相似文献

本文介绍了 Tofu,这是一个跨多个 GPU 设备划分非常大的 DNN 模型以减少每个 GPU 内存占用的系统。 Tofu 旨在对 MXNet 和 TensorFlow 等平台使用的细粒度张量运算符的数据流图进行分区。为了自动划分每个运算符,我们建议用受 Halide 启发的简单语言来描述运算符的语义。为了最佳地划分数据流图中的不同运算符,Tofu 使用递归搜索算法来最小化总通信成本。我们在 8-GPU 机器上的实验表明,Tofu 能够训练非常大的 CNN 和 RNN 模型。与训练超大型模型的替代方法相比,它还实现了 25% - 400% 的加速。
This paper presents Tofu, a system that partitions very large DNN models across multiple GPU devices to reduce per-GPU memory footprint. Tofu is designed to partition a dataflow graph of fine-grained tensor operators used by platforms like MXNet and TensorFlow. In order to automatically partition each operator, we propose to describe the semantics of an operator in a simple language inspired by Halide. To optimally partition different operators in a dataflow graph, Tofu uses a recursive search algorithm that minimizes the total communication cost. Our experiments on an 8-GPU machine show that Tofu enables the training of very large CNN and RNN models. It also achieves 25% - 400% speedup over alternative approaches to train very large models.