课题基金 / 基金详情

FaT HaMM: Fault Tolerant Hardware from Malleable Microarchitectures

FaT HaMM: Fault Tolerant Hardware from Malleable Microarchitectures
FaT HaMM:可延展微架构的容错硬件
批准号:
2283663
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

相关文献

中文摘要
翻译
主题:制造未来研究领域:人工智能技术现代电子容错建立在冗余的基础上。可演化硬件的基础研究强调了构建在灵活性基础上的容错系统的能力;如果发生关键故障,这些系统使用现场可编程门阵列(FPGA)和遗传算法来重写硬件配置。该提案概述了通过部署自动硬件重新校准来开发容错处理器的工作计划。将使用两种方法来缓解可演化硬件的历史上糟糕的可伸缩性。首先,完全可重构性将只在小规模上使用;通过在传统的不灵活的处理器中嵌入完全灵活的硬件的小区域。通过用构建在这种“可伸缩的”微体系结构上的对应组件替换处理器中的许多体系结构组件,我的目标是将可演化硬件系统的故障恢复功能构建为容易出现故障的体系结构组件。其次,灵活的区域将被分割成单独的区域,每个区域都是硅片区域,这样就可以精确地定位故障,并且只重新配置受影响的区域。这些步骤应该会将搜索空间减少到可管理的大小。可延展的微体系结构组件将包括类似于FPGA的可配置硬件区域的网状结构,以及在检测到故障时定位错误区域并重新配置它们的能力。这种实用的灵活性将分两部分开发:通过设计调整的机器学习算法来与遗传算法竞争或改进,以及通过使用可伸缩的微体系结构组件设计新的处理器体系结构。具体地说,这项研究计划的目标是:1.利用现代机器学习方法来开发新的自动化硬件设计方法;这些方法应该根据系统的规模而量身定做,并能够处理非平凡的硬件设计。通过工业伙伴关系改进现代处理器的已知故障模型。将现场可编程门阵列技术嵌入到传统的处理器设计中,创造出一种延展性强的硬件处理器,作为新的自动化硬件设计基板,具有可配置的灵活性。创建一个详尽的测试框架,探索所设计的延展性硬件系统的容错性和可伸缩性。机器学习研究的一个初步方向是探索深度学习作为硬件逻辑设计机制的潜力。即研究将深度学习的主动剪枝神经网络转换为逻辑门网络的可能性。如上所述,自动化硬件设计的核心问题之一是系统可伸缩性。为了解决这一问题,不是试图自动化整个处理器的设计,而是芯片的灵活部分将是独立的和可管理的大小;机器学习算法将被设计在可伸缩性的前沿。通过本研究的结论,希望能够构建出一种实用的故障恢复机制,与目前最先进的基于冗余的方法相抗衡。这些处理器将有能力重写有问题、容易出错的硬件,以从关键性能故障中恢复。理论上可延展的处理器可以预装配置,开箱即用,性能与传统芯片相同。然而,一旦检测到故障组件,搜索过程将开始寻找减少故障影响的替代硬件设计,或者完全否定故障。这将导致处理器能够进行有限的自我修复。
英文摘要
Theme: Manufacturing the FutureResearch Area: Artifical Inteligence TechnologiesModern fault tolerance in electronics is built on redundancy. Foundational research in evolvable hardware has highlighted the capacity for fault tolerant systems built on flexibility; should a critical fault occur, these systems use Field Programmable Gate Arrays (FPGAs) and genetic algorithms to rewrite hardware configurations. This proposal outlines a scheme of work to develop fault resistant processors by deploying automated hardware recalibration. Two approaches will be used to mitigate the historically poor scaling of evolvable hardware. Firstly, full reconfigurability will only be used on a small scale; by embedding small regions of fully flexible hardware within a conventional inflexible processor. By replacing a number of architectural components within a processor, with counterparts built on this "malleable" microarchitecture, I aim to build the fault recovery capabilities of evolvable hardware systems into fault prone architectural components. Secondly, the flexible region will be segmented into separate zones, each an area of silicon, so that the fault can be pinpointed and only the affected region reconfigured. These steps should reduce the search space to a manageable size.The malleable microarchitectural components will consist of a mesh of FPGA-like configurable areas of hardware, and the capacity to pinpoint erroneous areas and reconfigure them, if a fault is detected. This practical flexibility will be developed in two parts; by devising tuned machine learning algorithms to compete with, or improve on, genetic algorithms, and by designing a new processor architecture using malleable microarchitectural components.Specifically, this programme of research aims to:1. Exploit modern machine learning methods to develop new automated approaches to hardware design; these approaches should be tailored to system scaling and be capable of tackling non-trivial hardware designs.2. Improve known fault models for modern processors through industrial partnerships.3. Embed FPGA technology within conventional processor designs to create a malleable hardware processor to act as a new automated hardware design substrate, with configurable degrees of flexibility.4. Create an exhaustive testing framework and explore the fault tolerance and scaling properties of the designed malleable hardware system.An initial direction for the machine learning research would explore the potential of deep learning as a mechanism of hardware logic design. Namely looking into the possibility of converting deep-learned aggressively pruned neural nets into networks of logic gates.As mentioned, one of the central problems in automated hardware design is system scaling. To tackle this; rather than try to automate the design of an entire processor, the flexible portions of the chip will be self-contained and of a manageable size; and the machine learning algorithms will be designed with scalability at the forefront.By the conclusion of this research, it is hoped that a practical fault recovery mechanism can be constructed, rivalling the current cutting edge redundancy-based methods. These processors will have the capacity to rewrite problematic fault-prone hardware to recover from performance-critical faults. Theoretical malleable processors could ship with a configuration preloaded, and out of the box they will perform identically to a conventional chip. However, upon the detection of a faulty component, a search procedure will begin to look for alternate hardware designs which reduce the impact of the fault, or negate it completely. This will result in a processor capable of a limited amount of self-healing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文