Device Placement Optimization with Reinforcement Learning

Device Placement Optimization with Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2017-06
期刊:
--
影响因子:
--
通讯作者:
Azalia Mirhoseini;Hieu Pham;Quoc V. Le;Benoit Steiner;Rasmus Larsen;Yuefeng Zhou;Naveen Kumar;
Azalia Mirhoseini;Hieu Pham;Quoc V. Le;Benoit Steiner;Rasmus Larsen;Yuefeng Zhou;Naveen Kumar;
中科院分区:
其他
文献类型:
--
作者:
Azalia Mirhoseini;Hieu Pham;Quoc V. Le;Benoit Steiner;Rasmus Larsen;Yuefeng Zhou;Naveen Kumar;

文献摘要

被引文献

相似文献

在过去的几年里,神经网络的训练和推理的规模和计算需求都在增长。目前,解决这些需求的一种常见方法是使用混合了CPU和GPU等硬件设备的异构分布式环境。重要的是,将部分神经模型放置在设备上的决定通常是由人类专家基于简单的推理和直觉做出的。在本文中,我们提出了一种学习优化TensorFlow计算图的设备放置的方法。我们方法的关键是使用序列到序列模型来预测TensorFlow图中的哪些操作子集应该在哪些可用设备上运行。然后将预测放置的执行时间用作奖励信号以优化序列到序列模型的参数。我们的主要结果是,在Inception-V3上用于ImageNet分类,在RNN LSTM上用于语言建模和神经机器翻译,我们的模型发现了非平凡的设备放置,其性能优于手工制作的算法和传统的算法方法。
The past few years have witnessed a growth in size and computational requirements for training and inference with neural networks. Currently, a common approach to address these requirements is to use a heterogeneous distributed environment with a mixture of hardware devices such as CPUs and GPUs. Importantly, the decision of placing parts of the neural models on devices is often made by human experts based on simple heuristics and intuitions. In this paper, we propose a method which learns to optimize device placement for TensorFlow computational graphs. Key to our method is the use of a sequence-to-sequence model to predict which subsets of operations in a TensorFlow graph should run on which of the available devices. The execution time of the predicted placements is then used as the reward signal to optimize the parameters of the sequence-to-sequence model. Our main result is that on Inception-V3 for ImageNet classification, and on RNN LSTM, for language modeling and neural machine translation, our model finds non-trivial device placements that outperform hand-crafted heuristics and traditional algorithmic methods.