Fault-Tolerant Computing for Machine Learning Applications
Fault-Tolerant Computing for Machine Learning Applications
批准号:
RGPIN-2020-06884
负责人:
Nicolici, Nicola
金额:
$2.4万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31
中文摘要
机器学习已经成功地应用于消费者应用,如图像/语音识别。近年来,人们对在自主系统中采用机器学习越来越感兴趣。为了实现完全自主的目标,终身机器学习范式不仅必须成熟,而且必须保证其在可降解硬件上的运行。综上所述,这个项目的目的是研究如何在现场出现硬件故障的情况下,以一种优雅的可降解方式适应机器学习模型。在本研究计划的早期阶段,有必要了解故障硬件与机器学习应用程序交互的独特之处。特别是,除了一组顺序冗余故障(即在数字硬件的任何可达状态下无法激发的故障)之外,机器学习应用程序预计还会产生一大批应用程序冗余故障,即在给定的一组应用程序约束(例如,推理阶段的模型参数)下无法影响可观察输出的故障。此外,由于大多数机器学习应用程序在预测精度上都有用户可接受的损失,因此了解哪种类型的硬件故障在预测精度上产生可容忍的损失和不可容忍的损失同样重要。随后,我们将专注于开发针对机器学习工作负载特征的现场测试、诊断和容错的新方法。例如,在机器学习应用的推理阶段,如果硬件乘法器的其中一个操作数是常数,那么许多乘法器的内部网络将无法被观察到;因此,各个网络上的硬件故障是可以容忍的。这个简单的观察可以导致深入研究如何在机器学习硬件中存在的大量乘法器块上重新映射/重新调度节点/操作,以容忍一组已知故障。另外,也值得研究如何更新机器学习模型的参数以绕过故障,同时保证预测精度的可容忍损失。另一方面,在用于自主系统的强化学习环境中,需要确保即使存在硬件故障也能继续学习。这就提出了一个问题,即是否可以重新定义现有的在线学习算法,以确保模型参数不仅可以根据独特的操作环境进行调整,而且可以根据故障硬件进行调整。如上所述,机器学习工作负载为容错计算领域带来了新的维度。本研究计划的主要焦点是研究这些新的维度,并开发适用于广泛硬件架构的容错计算方法。
英文摘要
Machine learning has been successfully used in consumer applications, such as image/speech recognition. In recent years there has been a growing interest for adoption of machine learning in autonomous systems. In order to achieve the goal of full autonomy, not only lifelong machine learning paradigm must come of age, but also its guaranteed operation on degradable hardware is a necessity. Motivated by the above, the aim of this project is to investigate how machine learning models can be adapted in a gracefully degradable manner in the presence of hardware faults that arise in--field. In the early stages of this research program it will be necessary to understand what is unique about how faulty hardware interacts with machine learning applications. In particular, in addition to the set of sequentially--redundant faults, i.e., faults that cannot be excited in any reachable state of digital hardware, machine learning applications are expected to give rise to a large set of application--redundant faults, i.e., faults that cannot affect an observable output under a given set of application constraints, e.g., model parameters during the inference phase. Furthermore, since most machine learning applications have a user--acceptable loss in prediction accuracy, it is equally important to understand which types of hardware faults produce a tolerable vs an intolerable loss in prediction accuracy. Subsequently we will focus on developing novel methods for in--field test, diagnosis and fault tolerance that are specific to the characteristics of machine learning workloads. For example, during the inference phase of machine learning application, if one of the operands for a hardware multiplier is a constant then many of the multiplier's internal nets will not be observable; hence the hardware faults on the respective nets will be tolerated. This simple observation can lead to in--depth investigations on how to re-map/re-schedule nodes/operations on the large number of multiplier blocks present in machine learning hardware in order to tolerate a set of known faults. Alternatively it is also worth investigating how to update the parameters of a machine learning model in order to bypass the faults, while guaranteeing a tolerable loss in prediction accuracy. On another line of thought, in reinforcement learning environments used for autonomous systems, one needs to ensure that learning can continue despite the presence of hardware faults. This raises the question whether the existing on-line learning algorithms can be redefined in order to ensure that model parameters can be adjusted not only to the unique operating environment but also to the faulty hardware. As summarized above, machine learning workloads bring new dimensions to the field of fault--tolerant computing. It is the main focus of this research program to investigate these new dimensions and develop fault- tolerant computing methods adaptable to a broad spectrum of hardware architectures.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Fault-Tolerant Computing for Machine Learning Applications
-
批准号:RGPIN-2020-06884
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.4万
-
财政年份:2022
-
负责人:Nicolici, Nicola
-
依托单位:
Fault-Tolerant Computing for Machine Learning Applications
-
批准号:RGPIN-2020-06884
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.4万
-
财政年份:2020
-
负责人:Nicolici, Nicola
-
依托单位:
Systematic and Structural Methods for Post-Silicon Validation
-
批准号:RGPIN-2015-05312
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.7万
-
财政年份:2019
-
负责人:Nicolici, Nicola
-
依托单位:
Systematic and Structural Methods for Post-Silicon Validation
-
批准号:RGPIN-2015-05312
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.7万
-
财政年份:2018
-
负责人:Nicolici, Nicola
-
依托单位:
Systematic and Structural Methods for Post-Silicon Validation
-
批准号:478097-2015
-
项目类别:Discovery Grants Program - Accelerator Supplements
-
资助金额:$2.91万
-
财政年份:2017
-
负责人:Nicolici, Nicola
-
依托单位:
Systematic and Structural Methods for Post-Silicon Validation
-
批准号:RGPIN-2015-05312
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.7万
-
财政年份:2017
-
负责人:Nicolici, Nicola
-
依托单位:
Systematic and Structural Methods for Post-Silicon Validation
-
批准号:RGPIN-2015-05312
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.7万
-
财政年份:2016
-
负责人:Nicolici, Nicola
-
依托单位:
Systematic and Structural Methods for Post-Silicon Validation
-
批准号:RGPIN-2015-05312
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.7万
-
财政年份:2015
-
负责人:Nicolici, Nicola
-
依托单位:
Systematic and Structural Methods for Post-Silicon Validation
-
批准号:478097-2015
-
项目类别:Discovery Grants Program - Accelerator Supplements
-
资助金额:$2.91万
-
财政年份:2015
-
负责人:Nicolici, Nicola
-
依托单位:
Hardware accelerators for biomedical applications
-
批准号:239003-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.4万
-
财政年份:2014
-
负责人:Nicolici, Nicola
-
依托单位:
Hardware accelerators for biomedical applications
-
批准号:239003-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.4万
-
财政年份:2013
-
负责人:Nicolici, Nicola
-
依托单位:
Hardware accelerators for biomedical applications
-
批准号:239003-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.4万
-
财政年份:2012
-
负责人:Nicolici, Nicola
-
依托单位:
Hardware accelerators for biomedical applications
-
批准号:239003-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.4万
-
财政年份:2011
-
负责人:Nicolici, Nicola
-
依托单位:
Hardware accelerators for biomedical applications
-
批准号:239003-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.4万
-
财政年份:2010
-
负责人:Nicolici, Nicola
-
依托单位:
Embedded architectures for silicon debug, test and diagnosis
-
批准号:239003-2005
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2009
-
负责人:Nicolici, Nicola
-
依托单位:
Embedded architectures for silicon debug, test and diagnosis
-
批准号:239003-2005
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2008
-
负责人:Nicolici, Nicola
-
依托单位:
High quality and low cost system-on-a-chip test
-
批准号:326894-2005
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$2.33万
-
财政年份:2008
-
负责人:Nicolici, Nicola
-
依托单位:
High quality and low cost system-on-a-chip test
-
批准号:326894-2005
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$2.33万
-
财政年份:2006
-
负责人:Nicolici, Nicola
-
依托单位:
Embedded architectures for silicon debug, test and diagnosis
-
批准号:239003-2005
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2006
-
负责人:Nicolici, Nicola
-
依托单位:
Embedded architectures for silicon debug, test and diagnosis
-
批准号:239003-2005
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2005
-
负责人:Nicolici, Nicola
-
依托单位:
海外基金