Fine-grained parallelization of the Car-Parrinello ab initio molecular dynamics method on the IBM Blue Gene/L supercomputer

Fine-grained parallelization of the Car-Parrinello ab initio molecular dynamics method on the IBM Blue Gene/L supercomputer
复制标题

DOI:
10.1147/rd.521.0159
复制
发表时间:
2008-01-01
影响因子:
1.3
通讯作者:
Martyna, G. J.
Martyna, G. J.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Bohm, E.;Bhatele, A.;Martyna, G. J.

文献摘要

被引文献

相似文献

重要的科学问题可以通过基于从头计算的分子建模方法来处理,其中原子力来自明确考虑电子的能量函数。Car-Parrinello从头算分子动力学(CPAIMD)方法被广泛用于研究包含10到103个原子的小系统。然而,CPAIMD的影响一直是有限的,直到最近,因为固有的困难,以缩放技术超过处理器数量约等于电子状态的数量。CPAIMD计算涉及大量相互依赖的阶段,具有高处理器间通信开销。这些阶段需要评估各种变换和非方阵乘法,这些变换和非方阵乘法在高效并行化时需要大量的处理器间数据移动。使用Charm++并行编程语言和运行时系统,阶段被离散化为大量的虚拟处理器,这些虚拟处理器又被灵活地映射到物理处理器上,从而允许工作的交错。为了证明细粒度的并行性,采用了IBM Blue Gene/L(TM)系统特定的优化来将CPAIMD方法缩放到由24至768个原子(32至1,024个电子状态)组成的小系统中的电子状态数的至少30倍。研究的最大系统在整个机器(20,480个节点)上扩展良好。
Important scientific problems can be treated via ab initio-based molecular modeling approaches, wherein atomic forces are derived from an energy function that explicitly considers the electrons. The Car-Parrinello ab initio molecular dynamics (CPAIMD) method is widely used to study small systems containing on the order of 10 to 103 atoms. However, the impact of CPAIMD has been limited until recently because of difficulties inherent to scaling the technique beyond processor numbers about equal to the number of electronic states. CPAIMD computations involve a large number of interdependent phases with high interprocessor communication overhead. These phases require the evaluation of various transforms and non-square matrix multiplications that require large interprocessor data movement when efficiently parallelized. Using the Charm++ parallel programming language and runtime system, the phases are discretized into a large number of virtual processors, which are, in turn, mapped flexibly onto physical processors, thereby allowing interleaving of work. Algorithmic and IBM Blue Gene/L (TM) system-specfic optimizations are employed to scale the CPAIMD method to at least 30 times the number of electronic states in small systems consisting of 24 to 768 atoms (32 to 1,024 electronic states) in order to demonstrate fine-grained parallelism. The largest systems studied scaled well across the entire machine (20,480 nodes).