Colmena: Scalable Machine-Learning-Based Steering of Ensemble Simulations for High Performance Computing

Colmena: Scalable Machine-Learning-Based Steering of Ensemble Simulations for High Performance Computing
复制标题

DOI:
10.1109/mlhpc54614.2021.00007
复制
发表时间:
2021-10
期刊:
2021 IEEE/ACM Workshop on Machine Learning in High Performance Computing Environments (MLHPC)
影响因子:
--
通讯作者:
Logan T. Ward;G. Sivaraman;J. G. Pauloski;Y. Babuji;Ryan Chard;Naveen K. Dandu;P. Redfern;R. Assary;K. Chard;L. Curtiss;R. Thakur;Ian T. Foster
Logan T. Ward;G. Sivaraman;J. G. Pauloski;Y. Babuji;Ryan Chard;Naveen K. Dandu;P. Redfern;R. Assary;K. Chard;L. Curtiss;R. Thakur;Ian T. Foster
中科院分区:
其他
文献类型:
--
作者:
Logan T. Ward;G. Sivaraman;J. G. Pauloski;Y. Babuji;Ryan Chard;Naveen K. Dandu;P. Redfern;R. Assary;K. Chard;L. Curtiss;R. Thakur;Ian T. Foster

文献摘要

被引文献

相似文献

使用实验设计方法来选择最佳的模拟来执行,可以大大加快涉及模拟集成的科学应用。使用机器学习(ML)来创建模拟代理模型的方法显示出引导集合的特别前景,但由于需要协调模拟和学习任务的动态混合,因此部署具有挑战性。我们介绍了Colmena,这是一个开源的Python框架,允许用户通过提供单个任务的实现以及用于选择何时执行哪些任务的逻辑来引导活动。Colmena处理任务分派、结果整理、ML模型调用和ML模型(重新)训练,使用Parsl在HPC系统上执行任务。我们描述了设计的Colmena和说明其功能,将其应用到电解质设计,在那里它既可扩展到65536 CPU和加速发现率为高性能分子的100倍以上无指导的搜索。
Scientific applications that involve simulation ensembles can be accelerated greatly by using experiment design methods to select the best simulations to perform. Methods that use machine learning (ML) to create proxy models of simulations show particular promise for guiding ensembles but are challenging to deploy because of the need to coordinate dynamic mixes of simulation and learning tasks. We present Colmena, an open-source Python framework that allows users to steer campaigns by providing just the implementations of individual tasks plus the logic used to choose which tasks to execute when. Colmena handles task dispatch, results collation, ML model invocation, and ML model (re)training, using Parsl to execute tasks on HPC systems. We describe the design of Colmena and illustrate its capabilities by applying it to electrolyte design, where it both scales to 65536 CPUs and accelerates the discovery rate for high-performance molecules by a factor of 100 over unguided searches.