Operating Liquid-Cooled Large-Scale Systems: Long-Term Monitoring, Reliability Analysis, and Efficiency Measures

Operating Liquid-Cooled Large-Scale Systems: Long-Term Monitoring, Reliability Analysis, and Efficiency Measures
复制标题

DOI:
10.1109/hpca51647.2021.00078
复制
发表时间:
2021-02
期刊:
2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Rohan Basu Roy;Tirthak Patel;R. Kettimuthu;W. Allcock;Paul M. Rich;Adam Scovel;Devesh Tiwari
Rohan Basu Roy;Tirthak Patel;R. Kettimuthu;W. Allcock;Paul M. Rich;Adam Scovel;Devesh Tiwari
中科院分区:
其他
文献类型:
--
作者:
Rohan Basu Roy;Tirthak Patel;R. Kettimuthu;W. Allcock;Paul M. Rich;Adam Scovel;Devesh Tiwari

文献摘要

相似文献

由于其能源效率,过去十年液体冷却的使用有所增加。虽然之前的许多工作有助于在改善数据中心冷却方面取得进展,但其中绝大多数都是在短时间内对小型系统进行研究。计算机系统和 HPC 社区缺乏长期研究来强调运营液冷大型数据中心的挑战和解决方案。我们在六年的时间里对千万亿级超级计算机 Mira 进行了首次详细表征。该研究通过对环境指标的系统监测来实现,并讨论了新的研究途径,包括冷却剂监测故障。
The past decade has seen a rise in the use of liquid cooling due to its energy efficiency. While many previous works have helped make progress toward improving data center cooling, a vast majority of them perform studies on a small system over a short span. The computer systems and HPC community lacks a long-term study highlighting the challenges and solutions in operating a liquid-cooled large-scale data center. We conduct the first detailed characterization of a petascale supercomputer, Mira, over a span of six years. The study is enabled by systematic monitoring of the environmental metrics, and discusses new research avenues, including coolant monitor failures.