Scientific Highlight of the Month

Scientific Highlight of the Month
复制标题

DOI:
--
复制
发表时间:
--
期刊:
--
影响因子:
--
通讯作者:
A. Marek;V. Blum;R. Johanni;V. Havu;Bruno Lang;T. Auckenthaler;A. Heinecke;H. Bungartz;H. Lederer
A. Marek;V. Blum;R. Johanni;V. Havu;Bruno Lang;T. Auckenthaler;A. Heinecke;H. Bungartz;H. Lederer
中科院分区:
其他
文献类型:
--
作者:
A. Marek;V. Blum;R. Johanni;V. Havu;Bruno Lang;T. Auckenthaler;A. Heinecke;H. Bungartz;H. Lederer

文献摘要

被引文献

相似文献

获得大型矩阵的特征值和特征向量是电子结构理论和许多其他计算科学领域的关键问题。计算工作正式缩放为O(n 3),其大小为已研究的问题N,因此通常定义了实际计算无法克服的系统尺寸限制。在许多情况下,不仅需要占可能特征值/特征向量对的一小部分,因此仅关注少数特征值的迭代解决方案策略变得无效。同样,完全规避本特征值解决方案并不总是可取的或实用的。我们在这里回顾了有关密集特征值求解器的一些当前发展,然后专注于ELPA库,该库有助于对对称和遗传学特征值问题的有效代数解决方案,这些问题分别为具有实用值和复杂值的矩阵,分别为密集的矩阵。并行计算机平台。 ELPA依赖于Scalapack库的矩阵布局,解决标准和广义特征值问题,但用自己的子例程替换所有实际的平行解决方案步骤。最关键的时间步骤是将基质的矩阵还原为三角形形式和特征向量的相应反射变形。 ELPA既提供了一步的三角法(连续的住户转换),又提供了两步转换,尤其是针对较大矩阵和大量CPU核心的效率更高。 ELPA基于MPI标准,还提供了早期的混合MPI-OPENMPI实现。对于当前的高性能计算机架构(例如Cray或Intel/Infiniband),证明了在电子结构理论中出现的问题尺寸的10,000个CPU核心的可伸缩性。对于尺寸为260,000的矩阵,在蓝格烯/p上显示了高达295,000个CPU核心的可伸缩性。
Obtaining the eigenvalues and eigenvectors of large matrices is a key problem in electronic structure theory and many other areas of computational science. The computational effort formally scales as O(N 3) with the size of the investigated problem, N , and thus often defines the system size limit that practical calculations cannot overcome. In many cases, more than just a small fraction of the possible eigenvalue/eigenvector pairs is needed, so that iterative solution strategies that focus only on few eigenvalues become ineffective. Likewise, it is not always desirable or practical to circumvent the eigenvalue solution entirely. We here review some current developments regarding dense eigenvalue solvers and then focus on the ELPA library, which facilitates the efficient algebraic solution of symmetric and Hermitian eigen-value problems for dense matrices that have real-valued and complex-valued matrix entries, respectively, on parallel computer platforms. ELPA addresses standard as well as generalized eigenvalue problems, relying on the well documented matrix layout of the ScaLAPACK library but replacing all actual parallel solution steps with subroutines of its own. The most time-critical step is the reduction of the matrix to tridiagonal form and the corresponding backtransformation of the eigenvectors. ELPA offers both a one-step tridiagonalization (successive Householder transformations) and a two-step transformation that is more efficient especially towards larger matrices and larger numbers of CPU cores. ELPA is based on the MPI standard, with an early hybrid MPI-OpenMPI implementation available as well. Scalability beyond 10,000 CPU cores for problem sizes arising in the electronic structure theory is demonstrated for current high-performance computer architectures such as Cray or Intel/Infiniband. For a matrix of dimension 260,000, scalability up to 295,000 CPU cores has been shown on BlueGene/P.