Runtime Concurrency Control and Operation Scheduling for High Performance Neural Network Training

Runtime Concurrency Control and Operation Scheduling for High Performance Neural Network Training
复制标题

DOI:
10.1109/ipdps.2019.00029
复制
发表时间:
2018-10
期刊:
2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Jiawen Liu;Dong Li;Gokcen Kestor;J. Vetter
Jiawen Liu;Dong Li;Gokcen Kestor;J. Vetter
中科院分区:
其他
文献类型:
--
作者:
Jiawen Liu;Dong Li;Gokcen Kestor;J. Vetter

文献摘要

相似文献

训练神经网络(NN)通常使用TensorFlow和Caffe2等机器学习框架。这些框架采用数据流模型,其中神经网络训练被建模为由一组节点组成的有向图。神经网络训练中的操作通常由框架作为原语实现,并表示为数据流图中的节点。在基于数据流的机器学习框架中训练神经网络模型涉及大量的细粒度操作,这些操作呈现出不同的内存访问模式和计算强度。管理和调度这些操作具有挑战性,因为我们必须决定运行每个操作的线程数量(并发控制),并调度这些操作以获得良好的硬件利用率和系统吞吐量。在本文中,我们扩展了现有的运行时系统(TensorFlow运行时),以实现操作的自动并发控制和调度。我们探索性能建模来预测具有不同线程级并行性的操作的性能。我们的性能模型是高度精确和轻量级的。利用性能模型,我们的运行时系统采用了一组协同运行操作的调度策略,以提高硬件利用率和系统吞吐量。我们的运行时系统显示了显著的性能优势。与在TensorFlow中使用推荐配置进行并发控制和操作调度相比,我们的方法在四个神经网络模型上平均实现了36%的性能(执行时间)提升(高达49%),并且达到了接近用户手动获得的最优性能。
Training neural network (NN) often uses a machine learning framework such as TensorFlow and Caffe2. These frameworks employ a dataflow model where the NN training is modeled as a directed graph composed of a set of nodes. Operations in NN training are typically implemented by the frameworks as primitives and represented as nodes in the dataflow graph. Training NN models in a dataflow-based machine learning framework involves a large number of fine-grained operations whcih present diverse memory access patterns and computation intensity. Managing and scheduling those operations is challenging, because we have to decide the number of threads to run each operation (concurrency control) and schedule those operations for good hardware utilization and system throughput. In this paper, we extend an existing runtime system (the TensorFlow runtime) to enable automatic concurrency control and scheduling of operations. We explore performance modeling to predict the performance of operations with various thread-level parallelism. Our performance model is highly accurate and lightweight. Leveraging the performance model, our runtime system employs a set of scheduling strategies that co-run operations to improve hardware utilization and system throughput. Our runtime system demonstrates a significant performance benefit. Comparing with using the recommended configurations for concurrency control and operation scheduling in TensorFlow, our approach achieves 36% performance (execution time) improvement on average (up to 49%) for four neural network models, and achieves high performance close to the optimal one manually obtained by the user.