Leading Computational Methods on Scalar and Vector HEC Platforms

Leading Computational Methods on Scalar and Vector HEC Platforms
复制标题

标量和矢量 HEC 平台上的领先计算方法

DOI:
10.1109/sc.2005.41
复制
发表时间:
2005
期刊:
ACM/IEEE SC 2005 Conference (SC'05)
影响因子:
--
通讯作者:
Yoshinori Tsuda
Yoshinori Tsuda
中科院分区:
--
文献类型:
--
作者:
L. Oliker;J. Carter;M. Wehner;A. Canning;S. Ethier;A. Mirin;David Parks;P. Worley;S. Kitawaki;Yoshinori Tsuda

文献摘要

被引文献

相似文献

最近十年见证了超级高速缓存的微处理器的快速扩散,以构建高端计算(HEC)平台,这主要是因为它们的一般性,可扩展性和成本效益。但是,全面科学应用在常规超级计算机上的持续性能和峰值性能之间的差距不断增长,这已成为高性能计算的主要关注点,比峰值性能所暗示的更大的系统和可伸缩性需要更大的系统和应用可伸缩性,以实现所需的性能。最新一代的定制平行矢量系统有可能在其计算结构上有足够的规律性来解决此问题。在这项工作中,我们探讨了从四个领域提取的应用:大气建模(CAM),磁融合(GTC),等离子体物理(LBMHD3D)和材料科学(Paratec)。我们将基于矢量的Cray X1,Earth Simulator和新发行的NEC SX-8和Cray X1E的性能与使用IBM Power3,Intel Itanium2和AMD Opteron处理器的三个主要基于商品的Superscalar平台的性能进行了比较。我们的工作做出了一些重要的贡献:在高分辨率大气网上使用有限体积动力学核心的CAM模拟的第一个报道的矢量性能结果; GTC的新数据分解方案(首次)使TERAFLOP屏障的突破;引入了新的三维晶格Boltzmann Magneto-Hydrodnalnanic实现,用于研究血浆湍流的发作演化,在4800 ES基于promodity的Superscalar平台上实现了超过26Tflop/s的发作,利用IBM Power 3使用现代平行矢量系统:Cray X1,地球模拟器(ES)和NEC SX-8。此外,我们检查了CAM在最近发行的Cray X1E上的性能。我们的研究团队是第一个在地球模拟器中心进行绩效评估研究的国际小组。远程ES访问不可用。我们的工作以我们先前的努力为基础[16,17],并做出了一些重要的贡献:使用有限体积的动力学核心在高分辨率大气网格上使用有限体积的动力学核心进行了第一个报道的矢量性能结果; GTC的一种新的DatadeCoctsent(首次)使TERAFLOP屏障的突破;引入了新的三维晶格玻尔兹曼磁磁动力实现,用于研究4800 ES处理器上26Tflop/s的等离子体湍流的发作;以及迄今为止最大的Paratec细胞大小原子模拟。总体而言,结果表明,矢量体系结构在我们的应用程序套件中实现了前所未有的总体性能,这表明了现代平行矢量系统的巨大潜力。
The last decade has witnessed a rapid proliferation of superscalar cache-based microprocessors to build high-end computing (HEC) platforms, primarily because of their generality, scalability, and cost effectiveness. However, the growing gap between sustained and peak performance for full-scale scientific applications on conventional supercomputers has become a major concern in high performance computing, requiring significantly larger systems and application scalability than implied by peak performance in order to achieve desired performance. The latest generation of custom-built parallel vector systems have the potential to address this issue for numerical algorithms with sufficient regularity in their computational structure. In this work we explore applications drawn from four areas: atmospheric modeling (CAM), magnetic fusion (GTC), plasma physics (LBMHD3D), and material science (PARATEC). We compare performance of the vector-based Cray X1, Earth Simulator, and newly-released NEC SX-8 and Cray X1E, with performance of three leading commodity-based superscalar platforms utilizing the IBM Power3, Intel Itanium2, and AMD Opteron processors. Our work makes several significant contributions: the first reported vector performance results for CAM simulations utilizing a finite-volume dynamical core on a high-resolution atmospheric grid; a new data-decomposition scheme for GTC that (for the first time) enables a breakthrough of the Teraflop barrier; the introduction of a new three-dimensional Lattice Boltzmann magneto-hydrodynamic implementation used to study the onset evolution of plasma turbulence that achieves over 26Tflop/s on 4800 ES promodity-based superscalar platforms utilizing the IBM Power3, Intel Itanium2, and AMD Opteron processors, with modern parallel vector systems: the Cray X1, Earth Simulator (ES), and the NEC SX-8. Additionally, we examine performance of CAM on the recently-released Cray X1E. Our research team was the first international group to conduct a performance evaluation study at the Earth Simulator Center; remote ES access is not available. Our work builds on our previous efforts [16, 17] and makes several significant contributions: the first reported vector performance results for CAM simulations utilizing a finite-volume dynamical core on a high-resolution atmospheric grid; a new datadecomposition scheme for GTC that (for the first time) enables a breakthrough of the Teraflop barrier; the introduction of a new three-dimensional Lattice Boltzmann magneto-hydrodynamic implementation used to study the onset evolution of plasma turbulence that achieves over 26Tflop/s on 4800 ES processors; and the largest PARATEC cell size atomistic simulation to date. Overall, results show that the vector architectures attain unprecedented aggregate performance across our application suite, demonstrating the tremendous potential of modern parallel vector systems.