Energy efficiency vs. performance of the numerical solution of PDEs: An application study on a low-power ARM-based cluster

Energy efficiency vs. performance of the numerical solution of PDEs: An application study on a low-power ARM-based cluster
复制标题

偏微分方程数值解的能效与性能:基于 ARM 的低功耗集群的应用研究

DOI:
10.1016/j.jcp.2012.11.031
复制
发表时间:
2013
期刊:
J. Comput. Phys.
影响因子:
--
通讯作者:
Alex Ramírez
Alex Ramírez
中科院分区:
--
文献类型:
--
作者:
Dominik Göddeke;D. Komatitsch;M. Geveler;D. Ribbrock;Nikola Rajovic;Nikola Puzovic;Alex Ramírez

文献摘要

参考文献

被引文献

相似文献

功耗和能源效率正在成为大规模HPC设施设计和运行的关键方面,人们一致认为,未来的百亿亿次超级计算机将受到其功率要求的强烈约束。按照目前的电力成本,在HPC系统的整个生命周期内运行HPC系统的成本已经与初始部署成本相当。这些功耗限制,以及更节能的HPC平台可能对其他社会领域带来的好处,促使HPC研究界研究最初为嵌入式市场(尤其是移动的市场)开发的节能技术的使用。然而,较低的功率并不总是意味着较低的能耗,因为执行时间通常也会增加。为了实现有竞争力的性能,应用程序需要有效地利用大量的处理器。在本文中,我们将讨论应用程序如何有效地利用这类新的低功耗架构来实现具有竞争力的性能。我们评估他们是否可以从建筑应该实现的提高能源效率中受益。我们考虑的应用程序涵盖三个不同类别的偏微分方程的数值求解方法,即一个低阶有限元多重网格求解器的巨大稀疏线性方程组,格子玻尔兹曼代码的流体模拟,和高阶谱元方法的声波或地震波传播建模。我们评估了96个ARM Cortex-A9双核处理器集群的弱可扩展性和强可扩展性,并证明了基于ARM的集群在执行三个应用程序时,与基于x86的参考机相比,在解决方案的能量方面可以更有效。
Power consumption and energy efficiency are becoming critical aspects in the design and operation of large scale HPC facilities, and it is unanimously recognised that future exascale supercomputers will be strongly constrained by their power requirements. At current electricity costs, operating an HPC system over its lifetime can already be on par with the initial deployment cost. These power consumption constraints, and the benefits a more energy-efficient HPC platform may have on other societal areas, have motivated the HPC research community to investigate the use of energy-efficient technologies originally developed for the embedded and especially mobile markets. However, lower power does not always mean lower energy consumption, since execution time often also increases. In order to achieve competitive performance, applications then need to efficiently exploit a larger number of processors. In this article, we discuss how applications can efficiently exploit this new class of low-power architectures to achieve competitive performance. We evaluate if they can benefit from the increased energy efficiency that the architecture is supposed to achieve. The applications that we consider cover three different classes of numerical solution methods for partial differential equations, namely a low-order finite element multigrid solver for huge sparse linear systems of equations, a Lattice-Boltzmann code for fluid simulation, and a high-order spectral element method for acoustic or seismic wave propagation modelling. We evaluate weak and strong scalability on a cluster of 96 ARM Cortex-A9 dual-core processors and demonstrate that the ARM-based cluster can be more efficient in terms of energy to solution when executing the three applications compared to an x86-based reference machine.
DOI: 10.1177/1094342010391989
发表时间: 2011-02-01
影响因子: 3.1
作者:
Dongarra, Jack;Beckman, Pete;Yelick, Kathy
通讯作者: Yelick, Kathy
DOI: 10.1016/j.jcp.2010.06.024
发表时间: 2010-10-01
影响因子: 4.1
作者:
Komatitsch, Dimitri;Erlebacher, Gordon;Michea, David
通讯作者: Michea, David