An experimental exploration of self-aware systems for exascale architectures

An experimental exploration of self-aware systems for exascale architectures
复制标题

DOI:
--
复制
发表时间:
2016
期刊:
--
影响因子:
--
通讯作者:
A. Landwehr
A. Landwehr
中科院分区:
其他
文献类型:
--
作者:
A. Landwehr

文献摘要

被引文献

相似文献

高性能系统正在发展到性能不再是唯一相关标准的地步。当前的执行和资源管理范例已不足以确保正确性和性能。电力需求目前正在推动 HPC 系统的协同设计,这反过来又为如何表达对越来越稀缺资源的需求以及如何管理这些资源的根本改变奠定了基础。因此,系统需要在性能、能量和弹性方面变得更加内省和自我意识。为此,本文探讨了对实现自省至关重要的主要硬件要求、自省系统软件所需的接口类型和信息,提供了基于当前趋势的百亿亿次架构的抽象表示,并实现了具有内置温度和电源管理功能的百亿亿次模拟框架。通过这个框架,我们证明局部自适应策略对于百亿亿级系统来说是不够的,而是需要协调的分层自适应策略,以便有效地适应和减轻由数千个独立核心组成的系统内的振荡。
High-performance systems are evolving to a point where performance is no longer the sole relevant criterion anymore. Current execution and resource management paradigms are no longer sufficient to ensure correctness and performance. Power requirements are presently driving the co-design of HPC systems, which in turn sets the course for a radical change in how to express the need for scarcer and scarcer resources, as well as, how to manage them. As a result, systems will need to become more introspective and self-aware with respect to performance, energy, and resiliency. To this end, this thesis explores the major hardware requirements that are central to enabling introspection, the types of interfaces and information that will be needed for introspective system software, provides an abstract representation of exascale architectures based on current trends, and implements an exascale simulation framework with built in temperature and power management capabilities. Through this framework, we demonstrate that localized adaptive policies are not sufficient for exascale systems and that instead coordinated hierarchical adaptive policies are need in order to effectively adapt and mitigate oscillation within systems consisting of thousands of independent cores.