Towards a Graceful Degradable Multicore-System by Hierarchical Handling of Hard Errors

Towards a Graceful Degradable Multicore-System by Hierarchical Handling of Hard Errors
复制标题

通过硬错误的分层处理实现优雅的可降级多核系统

DOI:
10.1109/pdp.2013.51
复制
发表时间:
2013
期刊:
2013 21st Euromicro International Conference on Parallel, Distributed, and Network-Based Processing
影响因子:
--
通讯作者:
H. Vierhaus
H. Vierhaus
中科院分区:
--
文献类型:
--
作者:
Sebastian Müller;Mario Schölzel;H. Vierhaus

文献摘要

被引文献

相似文献

我们提出了一个新的概念,用于处理永久性故障在一个静态调度的异构多核系统通过基于软件的自我重新配置。硬故障以分层的跨层方式处理,或者由每个核心本身本地处理,或者通过重新配置整个系统来全局处理。缺陷核的局部重新配置基于所执行的任务对核的当前故障状态的适配,使得从不使用缺陷组件。这种自适应是通过重新调度任务的程序代码来实现的。如果该本地重新配置失败,则修改任务到核的绑定。因为允许异构核,所以这可能需要重新调度绑定被改变的任务。估计这样的全球重新配置的运行时间。此外,它表明,支持全球重新配置的系统实现了相同的容错水平的系统与本地修复只,但减少了硬件开销。
We present a novel concept for handling permanent faults in a statically scheduled heterogeneous multi-core system by means of a software-based self-reconfiguration. Hard faults are handled in a hierarchical cross layer manner, either locally by each core itself, or globally by reconfiguring the full system. Local reconfiguration of a defect core is based on the adaptation of the executed task to the current fault state of the core, such that defect components are never used. This adaptation is achieved by a rescheduling of the program code of the task. If this local reconfiguration fails, then the binding of tasks to cores is modified. Because heterogeneous cores are allowed, this may require a rescheduling of the tasks whose binding is changed. Estimations for the runtime of such a global reconfiguration are presented. Moreover, it is shown that systems that support the global reconfiguration achieve the same fault tolerance level as systems with local repair only, but with reduced hardware overhead.