课题基金 / 基金详情

PPoSS: Planning: CP2: Towards Systems Correctness Checkability and Performance Predictability at Scale

PPoSS: Planning: CP2: Towards Systems Correctness Checkability and Performance Predictability at Scale
PPoSS:规划:CP2:实现大规模系统正确性可检查性和性能可预测性
批准号:
2028427
负责人:
Haryadi Gunawi
金额:
$24.8万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-01 至 2022-08-31

项目摘要

项目成果

Haryadi Gunawi的其他基金

相似基金

相关文献

中文摘要
翻译
作为当今许多应用程序和服务的关键后端,大规模分布式系统必须具有高可靠性。在过去几年中,该领域的部署规模惊人;众所周知,谷歌运行的集群每个都有数千台机器,苹果部署了超过10万台数据库机器,Netflix运行的是数十个数据库集群,每个集群有500个节点。这个云规模分布式系统的新时代产生了一类新的故障,即可伸缩性故障——这些故障的症状在大规模部署中显现出来,但在中小型部署中却不一定。CP2方案是为了解决极端规模下系统的正确性可检性和性能可预测性问题而提出的。具体来说,该项目将分析十几个大型系统中的500多个真实可伸缩性故障,开发一个单机规模检查框架,允许开发人员在一台或几台机器上测试大型分布式代码,并为现有和未来架构上的大规模作业的计算和I/ o性能可预测性提供基础。这些任务将推进传统硬件平台和新兴硬件平台上的调试、测试、学习和预测方法,并最终导致构建正确的开发方法。CP2项目将对多个学科产生影响,包括系统(云/数据中心系统可靠性)、编程语言/编译器(新的静态/动态分析技术)、架构(异构硬件的计算/存储预测)、算法(学习方法的使用)和高性能计算(高性能计算系统/应用程序的基准测试)。在社会效益方面,CP2项目解决了国家科学基金会2018-2022年战略计划中提到的最重要问题。更具体地说,社会越来越依赖于复杂的系统,这些系统是人类聪明才智的产物,包括由大型复杂软件组成的生态系统,这些软件在数千台机器上运行着数百万行代码。CP2将解决理解和预测这些系统行为的挑战。此外,随着社会对复杂系统的依赖日益增长,了解它们的稳健性并了解如何加强它们变得越来越重要。在教育方面,CP2项目提供独特的实践研究和尖端系统技术教育,学生将接受培训,在大量机器上操作软件,并分析其性能和正确性。CP2项目的成果将通过经典的出版媒介发布,通过开发大量开源的软件工件,最后通过与各种行业伙伴的合作来帮助塑造下一代大规模系统。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
As a critical backend for many of today's applications and services, large-scale distributed systems must be highly reliable. In the last couple of years the field witnessed a phenomenal scale of deployment; Google is known to run clusters with thousands of machines each, Apple deploys over 100,000 database machines, and Netflix runs tens of database clusters with 500 nodes each. This new era of cloud-scale distributed systems has given birth to a new class of faults, scalability faults---faults whose symptoms surface in large-scale deployments but not necessarily in small/medium-scale deployments. The CP2 project is proposed to solve the problem of correctness checkability and performance predictability of systems at extreme scale. Specifically the project will analyze over 500 real-world scalability faults in over a dozen large-scale systems, develop a single-machine scale-checking framework that allows developers to test large distributed code on one or a few machines, and provide groundwork for compute- and I/O-performance predictability of large-scale jobs on both existing and future architectures. These tasks will advance debugging, testing, learning, and prediction methods both on traditional hardware platforms and emerging ones and ultimately lead to correct-by-construction development methods. The CP2 project will have impact in multiple disciplines including systems (cloud/datacenter systems reliability), programming languages/compilers (new static/dynamic analysis techniques), architecture (compute/storage prediction for heterogeneous hardware), algorithms (the use of learning methods), and high-performance computing (benchmarking of HPC systems/applications).In terms of societal benefits, the CP2 project addresses paramount issues mentioned in the NSF Strategic Plan for 2018-2022. More specifically, society increasingly depends on complicated systems that are products of human ingenuity, including ecosystems of large and complex software with millions of lines of code running on thousands of machines. CP2 will address the challenges of understanding and predicting the behavior of such systems. Furthermore, as society’s reliance on complex systems grows, learning about their robustness and understanding how to strengthen them are of increasing importance. In terms of education, the CP2 project gives unique hands-on research and education with cutting-edge systems technology in which students will be trained to operate software on a large number of machines and analyze their performance and correctness. The results of the CP2 project will be released through the classic medium of publication, through the development of numerous software artifacts which will be open-sourced, and finally through collaboration with various industry partners to help shape the next generation of large-scale systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: PPoSS: LARGE: ScaleStuds: Foundations for Correctness Checkability and Performance Predictability of Systems at Scale
  • 批准号:
    2119184
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $312.5万
  • 财政年份:
    2021
  • 负责人:
    Haryadi Gunawi
  • 依托单位:
USENIX FAST 2017 NSF Student Travel Support
  • 批准号:
    1727380
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2017
  • 负责人:
    Haryadi Gunawi
  • 依托单位:
CSR: Medium:Combating Distributed Concurrency Bugs in Cloud Systems
  • 批准号:
    1563956
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $80.0万
  • 财政年份:
    2016
  • 负责人:
    Haryadi Gunawi
  • 依托单位:
CSR: Small: BreezeFS: File System Transformation for Cloud and Multistore Era
  • 批准号:
    1526304
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.8万
  • 财政年份:
    2015
  • 负责人:
    Haryadi Gunawi
  • 依托单位:
海外基金