CAREER: Towards Gray-Fault Tolerant Cloud through Harnessing and Enhancing System Observability
CAREER: Towards Gray-Fault Tolerant Cloud through Harnessing and Enhancing System Observability
批准号:
2317751
负责人:
Peng Huang
金额:
$60.95万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-03-15 至 2025-08-31
中文摘要
云系统是当今许多现有服务的关键基础设施。确保云软件不间断地持续运行是至关重要的,也是具有挑战性的。几十年的研究已经发展出成熟的技术来检测和屏蔽分布式系统中的故障。但这些技术通常使用一个简单的模型,该模型假设系统组件要么工作,要么完全停止。然而,大量现实世界的云事件表明,生产云系统经常经历灰色故障-一种降级的操作模式,在这种模式下,系统组件似乎正在工作,但实际上严重受损。灰色故障不能用当前的解决方案有效地处理。该提案的总体目标是开发一种全面的方法来检测、定位和诊断生产云系统中的灰色故障。为了实现这一目标,提出了四项协同研究活动。具体地说,该项目对流行的分布式系统中的真实世界灰色故障案例进行研究,测量和表征现有系统的可观测性。然后,该项目设计了一种新的混合分析,该分析自动在整个系统堆栈中插入报告生成挂钩,以利用可观察性来检测灰色故障。为了找出罪魁祸首,该项目进一步提出了从收集的观测数据中推断因果关系的算法。最后,为了提高灰色故障的可观测性和在线诊断能力,本项目设计了一个运行时检测框架。灰色故障是云服务中断的常见原因,导致重大经济损失。该项目可以有效地提高我们对灰色故障的理解,帮助检测和调试灰色故障,以减少其对无处不在的云基础设施的影响。随着越来越多的细微故障模式,软件正朝着更加分布式的方向发展。可观察性、故障检测和本地化是这种范式转换的关键技能,但在现有课程中很少涉及。该项目通过课程开发和学生培训来解决这一教育差距。该项目还通过与非营利性组织Code in the School合作为当地高中生举办研讨会,展示云和系统故障的概念,向未被充分代表的巴尔的摩高中生推广计算机科学教育。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Cloud systems are the crucial infrastructure to many services existing today. Ensuring cloud software runs continuously without disruptions is both vital and challenging. Decades of research have developed mature techniques to detect and mask faults in distributed systems. But these techniques often use a simple model that assumes a system component either works or completely stops. Numerous real-world cloud incidents, however, suggest that production cloud systems frequently experience gray failures---a degraded operational mode in which a system component appears to be working but is in fact severely impaired. Gray failures cannot be effectively dealt with by current solutions. The overall objective of this proposal is to develop a holistic approach to detect, pinpoint and diagnose gray failures in production cloud systems. To realize the objective, four synergistic research activities are proposed. Specifically, the project conducts a study on real-world gray failure cases in popular distributed systems, measure and characterize the observability of existing systems. The project then designs a novel hybrid analysis that automatically inserts report-generation hooks across the whole systems stack to harness observability for detecting gray failures. To pinpoint the culprit component, this project further proposes algorithms to infer causality from the collected observations. Lastly, this project designs a runtime checking framework for increasing observability and online diagnosis of gray failures. Gray failures are a common cause of cloud service outages, resulting in significant financial loss. This project can effectively improve our understandings of gray failures and help detect and debug gray failures to reduce their impact on the ubiquitous cloud infrastructures. Software is moving to be more distributed with increasing subtle failure modes. Observability, fault detection, and localization are critical skills for this paradigm shift but are rarely covered in the existing curriculum. This project addresses this educational gap through curriculum development and student training. This project also promotes Computer Science education to underrepresented Baltimore high school students by organizing workshops in partnership with a non-profit organization, Code in the Schools, for local high school students to showcase cloud and system failure concepts.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1145/3600006.3613159
发表时间:
2023-10
期刊:
Proceedings of the 29th Symposium on Operating Systems Principles
影响因子:
--
作者:
[Yigong Hu;Gongqi Huang;Peng Huang]
通讯作者:
Yigong Hu;Gongqi Huang;Peng Huang
DOI:
10.1145/3626111.3628206
发表时间:
2023
期刊:
ACM
影响因子:
--
作者:
[Qiu, Yiming, Kon, Patrick Tser, Xing, Jiarong, Huang, Yibo, Liu, Hongyi, Wang, Xinyu, Huang, Peng, Chowdhury, Mosharaf, Chen, Ang]
通讯作者:
Chen, Ang
CNS Core: Small: Intelligent Fault Injection to Expose and Reproduce Production-Grade Bugs in Cloud Systems
-
批准号:2317698
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2023
-
负责人:Peng Huang
-
依托单位:
FMitF: Track I: Synthesizing Semantic Checkers for Runtime Verification of Production Distributed Systems
-
批准号:2318937
-
项目类别:Standard Grant
-
资助金额:$75.0万
-
财政年份:2023
-
负责人:Peng Huang
-
依托单位:
CNS Core: Small: Intelligent Fault Injection to Expose and Reproduce Production-Grade Bugs in Cloud Systems
-
批准号:2149664
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2021
-
负责人:Peng Huang
-
依托单位:
CAREER: Towards Gray-Fault Tolerant Cloud through Harnessing and Enhancing System Observability
-
批准号:1942794
-
项目类别:Continuing Grant
-
资助金额:$60.95万
-
财政年份:2020
-
负责人:Peng Huang
-
依托单位:
CRII: CSR: Toward Understanding and Automatically Detecting Specious Configuration in Large Systems
-
批准号:1755737
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2018
-
负责人:Peng Huang
-
依托单位:
海外基金