Low communication high performance ab initio density matrix renormalization group algorithms

Low communication high performance ab initio density matrix renormalization group algorithms
复制标题

DOI:
10.1063/5.0050902
复制
发表时间:
2021-06-14
影响因子:
4.4
通讯作者:
Chan, Garnet Kin-Lic
Chan, Garnet Kin-Lic
中科院分区:
化学2区
文献类型:
--
作者:
Zhai, Huanchen;Chan, Garnet Kin-Lic

文献摘要

被引文献

相似文献

最近,人们对在高性能计算平台上部署从头算密度矩阵重整化群(DMRG)计算感兴趣。在这里,我们介绍了一种传统的分布式存储器从头开始DMRG算法的重新表述,它将其与概念上更简单和有利的子哈密尔顿方法的和联系起来。从这个框架出发,我们进一步探索了并行策略的层次结构,包括(I)子哈密顿和上的并行,(Ii)站点上的并行,(Iii)正规和互补算子上的并行,(Iv)对称扇区上的并行,以及(V)密集矩阵乘法上的并行。我们描述了如何减少处理器负载不平衡和算法的通信开销,以达到更高的效率。我们说明了我们新的开源实现在最近的基准基态计算中的性能,苯在108个轨道和30个电子的轨道空间中,键维高达6000,以及一个具有76个轨道和113个电子的FEMO余因模型。观察到的从448个中央处理器核心到2800个中央处理单元核心的并行扩展几乎是理想的。由AIP出版公司独家授权出版。
There has been recent interest in the deployment of ab initio density matrix renormalization group (DMRG) computations on high performance computing platforms. Here, we introduce a reformulation of the conventional distributed memory ab initio DMRG algorithm that connects it to the conceptually simpler and advantageous sum of the sub-Hamiltonian approach. Starting from this framework, we further explore a hierarchy of parallelism strategies that includes (i) parallelism over the sum of sub-Hamiltonians, (ii) parallelism over sites, (iii) parallelism over normal and complementary operators, (iv) parallelism over symmetry sectors, and (v) parallelism within dense matrix multiplications. We describe how to reduce processor load imbalance and the communication cost of the algorithm to achieve higher efficiencies. We illustrate the performance of our new open-source implementation on a recent benchmark ground-state calculation of benzene in an orbital space of 108 orbitals and 30 electrons, with a bond dimension of up to 6000, and a model of the FeMo cofactor with 76 orbitals and 113 electrons. The observed parallel scaling from 448 to 2800 central processing unit cores is nearly ideal. Published under an exclusive license by AIP Publishing.