Understanding GPU errors on large-scale HPC systems and the implications for system design and operation

Understanding GPU errors on large-scale HPC systems and the implications for system design and operation
复制标题

DOI:
10.1109/hpca.2015.7056044
复制
发表时间:
2015-03
期刊:
2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Devesh Tiwari;Saurabh Gupta;James H. Rogers;Don E. Maxwell;P. Rech;Sudharshan S. Vazhkudai;Daniel Oliveira;Dave Londo;Nathan Debardeleben;P. Navaux;L. Carro;Arthur S. Bland
Devesh Tiwari;Saurabh Gupta;James H. Rogers;Don E. Maxwell;P. Rech;Sudharshan S. Vazhkudai;Daniel Oliveira;Dave Londo;Nathan Debardeleben;P. Navaux;L. Carro;Arthur S. Bland
中科院分区:
其他
文献类型:
--
作者:
Devesh Tiwari;Saurabh Gupta;James H. Rogers;Don E. Maxwell;P. Rech;Sudharshan S. Vazhkudai;Daniel Oliveira;Dave Londo;Nathan Debardeleben;P. Navaux;L. Carro;Arthur S. Bland

文献摘要

被引文献

相似文献

图形硬件性能的提高和可编程性的改进使得GPU能够从图形专用加速器发展为通用计算设备。Titan是2014年世界上第二快的开放科学超级计算机,由更多的dum 18,000 GPU组成,来自天体物理学,聚变,气候和燃烧等各个领域的科学家经常使用这些GPU来运行大规模模拟。不幸的是,虽然GPU的性能效率很好地理解,但它们在大规模计算系统中的弹性特性尚未得到充分评估。我们提出了一个详细的研究,以提供一个全面的了解GPU的错误在一个大规模的GPU支持的系统。我们的数据是从橡树岭领导计算设施的泰坦超级计算机和洛斯阿拉莫斯国家实验室的GPU集群收集的。我们还展示了在洛斯阿拉莫斯中子科学中心(LANSCE)和ISIS(Rutherford Appleron Laboratories,UK)进行的广泛中子束测试的结果,以测量不同代GPU的弹性。我们从我们的现场数据和中子束实验中提出了几项发现,并讨论了我们的结果对未来GPU架构师,当前和未来的HPC计算设施以及专注于GPU弹性的研究人员的影响。
Increase in graphics hardware performance and improvements in programmability has enabled GPUs to evolve from a graphics-specific accelerator to a general-purpose computing device. Titan, the world's second fastest supercomputer for open science in 2014, consists of more dum 18,000 GPUs that scientists from various domains such as astrophysics, fusion, climate, and combustion use routinely to run large-scale simulations. Unfortunately, while the performance efficiency of GPUs is well understood, their resilience characteristics in a large-scale computing system have not been fully evaluated. We present a detailed study to provide a thorough understanding of GPU errors on a large-scale GPU-enabled system. Our data was collected from the Titan supercomputer at the Oak Ridge Leadership Computing Facility and a GPU cluster at the Los Alamos National Laboratory. We also present results from our extensive neutron-beam tests, conducted at Los Alamos Neutron Science Center (LANSCE) and at ISIS (Rutherford Appleron Laboratories, UK), to measure the resilience of different generations of GPUs. We present several findings from our field data and neutron-beam experiments, and discuss the implications of our results for future GPU architects, current and future HPC computing facilities, and researchers focusing on GPU resilience.