Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models
复制标题

DOI:
10.1109/mlhpcai4s51975.2020.00013
复制
发表时间:
2020-07
期刊:
2020 IEEE/ACM Workshop on Machine Learning in High Performance Computing Environments (MLHPC) and Workshop on Artificial Intelligence and Machine Learning for Scientific Applications (AI4S)
影响因子:
--
通讯作者:
Sergio Botelho;Ameya Joshi;Biswajit Khara;S. Sarkar;C. Hegde;Santi S. Adavani;B. Ganapathysubramanian
Sergio Botelho;Ameya Joshi;Biswajit Khara;S. Sarkar;C. Hegde;Santi S. Adavani;B. Ganapathysubramanian
中科院分区:
其他
文献类型:
--
作者:
Sergio Botelho;Ameya Joshi;Biswajit Khara;S. Sarkar;C. Hegde;Santi S. Adavani;B. Ganapathysubramanian

文献摘要

相似文献

科学机器学习(SciML)的最新进展为训练解决复杂偏微分方程(PDE)的新型神经网络架构开辟了可能性。最近有几种(几乎无数据)方法成功地解决了PDE,例如深度前馈网络,生成网络和深度编码器-解码器网络。然而,这些方法的实际应用受到训练这些模型的困难的限制,特别是在大输出分辨率(≥ 1024 × 1024)下进行预测。在这里,我们报告了一个用于数据并行分布式深度学习的软件框架,它解决了在合理的时间内训练这些大型SciML模型以及分布存储需求的双重挑战。我们的框架提供了几个开箱即用的功能,包括(a)损失的完整性独立于进程的数量,(B)同步批量归一化,(c)分布式高阶优化方法。我们展示了这个框架在云和HPC集群上的出色可扩展性,并报告了带宽,网络拓扑结构和裸机与云之间的相互作用。我们部署这种方法来训练迄今为止不可能的生成模型,表明神经PDE求解器可以为实际应用进行可行的训练。我们还证明了分布式高阶优化方法比基于随机梯度的方法快2-3倍,并且在更高的批量大小下提供最小的收敛漂移。
Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥ 1024 × 1024).Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods.We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2–3 × faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.