Optimizing Excited-State Electronic-Structure Codes for Intel Knights Landing: A Case Study on the BerkeleyGW Software

Optimizing Excited-State Electronic-Structure Codes for Intel Knights Landing: A Case Study on the BerkeleyGW Software
复制标题

优化 Intel Knights Landing 的激发态电子结构代码:BerkeleyGW 软件案例研究

DOI:
--
复制
发表时间:
2016
期刊:
ISC Workshops
影响因子:
--
通讯作者:
S. Louie
S. Louie
中科院分区:
--
文献类型:
--
作者:
J. Deslippe;F. Jornada;Derek Vigil;Taylor A. Barnes;N. Wichmann;Karthik Raman;Ruchira Sasanka;S. Louie

文献摘要

被引文献

相似文献

我们分析并优化了在Xeon-Phi架构上使用BerkeleyGW [2,3]代码进行的计算。BerkeleyGW依赖于手动调整的关键内核以及BLAS和FFT库。我们描述的优化过程和实现的性能改进。我们讨论了分层并行化策略,以利用向量,线程和节点级并行。我们讨论的本地化的变化(包括缺乏L3缓存的后果)和有效地使用的封装高带宽内存。我们展示了骑士登陆的初步结果,包括一些优化前后的代码性能的屋顶研究。我们发现GW方法特别适合于众核架构,因为它能够在平面波分量、频带对和频率上利用大量的并行性。
We profile and optimize calculations performed with the BerkeleyGW [2, 3] code on the Xeon-Phi architecture. BerkeleyGW depends both on hand-tuned critical kernels as well as on BLAS and FFT libraries. We describe the optimization process and performance improvements achieved. We discuss a layered parallelization strategy to take advantage of vector, thread and node-level parallelism. We discuss locality changes (including the consequence of the lack of L3 cache) and effective use of the on-package high-bandwidth memory. We show preliminary results on Knights-Landing including a roofline study of code performance before and after a number of optimizations. We find that the GW method is particularly well-suited for many-core architectures due to the ability to exploit a large amount of parallelism over plane-wave components, band-pairs, and frequencies.