CosmoFlow: Using Deep Learning to Learn the Universe at Scale

CosmoFlow: Using Deep Learning to Learn the Universe at Scale
复制标题

DOI:
10.1109/sc.2018.00068
复制
发表时间:
2018-08
期刊:
SC18: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Amrita Mathuriya;D. Bard;P. Mendygral;Lawrence Meadows;James A. Arnemann;Lei Shao;Siyu He;Tuomas Kärnä;Diana Moise;S. Pennycook;K. Maschhoff;J. Sewall;Nalini Kumar;S. Ho;Michael F. Ringenburg;P. Prabhat;Victor W. Lee
Amrita Mathuriya;D. Bard;P. Mendygral;Lawrence Meadows;James A. Arnemann;Lei Shao;Siyu He;Tuomas Kärnä;Diana Moise;S. Pennycook;K. Maschhoff;J. Sewall;Nalini Kumar;S. Ho;Michael F. Ringenburg;P. Prabhat;Victor W. Lee
中科院分区:
其他
文献类型:
--
作者:
Amrita Mathuriya;D. Bard;P. Mendygral;Lawrence Meadows;James A. Arnemann;Lei Shao;Siyu He;Tuomas Kärnä;Diana Moise;S. Pennycook;K. Maschhoff;J. Sewall;Nalini Kumar;S. Ho;Michael F. Ringenburg;P. Prabhat;Victor W. Lee

文献摘要

被引文献

相似文献

深度学习是确定描述我们宇宙的物理模型的一个很有前途的工具。为了处理这个问题的大量计算成本,我们提出了CosmoFlow:一个建立在TensorFlow框架之上的高度可扩展的深度学习应用程序。CosmoFlow使用3D卷积和池原语的高效实现,以及对许多元素操作的线程改进,以提高英特尔®Xeon Phi™处理器的训练性能。我们还利用Cray PE机器学习插件来高效地扩展到多个节点。我们在Cori的8192个节点上展示了完全同步的数据并行训练,并行效率为77%,实现了3.5 Pflop/s的持续性能。据我们所知,这是TensorFlow框架在超级计算机规模上的第一次大规模科学应用,具有完全同步的训练。这些改进使我们能够处理大型三维暗物质分布,并以前所未有的精度预测宇宙参数ΩsubM/sub, σsub8/sub和nsubs/sub。
Deep learning is a promising tool to determine the physical model that describes our universe. To handle the considerable computational cost of this problem, we present CosmoFlow: a highly scalable deep learning application built on top of the TensorFlow framework. CosmoFlow uses efficient implementations of 3D convolution and pooling primitives, together with improvements in threading for many element-wise operations, to improve training performance on Intel® Xeon Phi™ processors. We also utilize the Cray PE Machine Learning Plugin for efficient scaling to multiple nodes. We demonstrate fully synchronous data-parallel training on 8192 nodes of Cori with 77% parallel efficiency, achieving 3.5 Pflop/s sustained performance. To our knowledge, this is the first large-scale science application of the TensorFlow framework at supercomputer scale with fully-synchronous training. These enhancements enable us to process large 3D dark matter distribution and predict the cosmological parameters ΩsubM/sub, σsub8/sub and nsubs/sub with unprecedented accuracy.