Collaborative Research: PPoSS: LARGE: ScaleStuds: Foundations for Correctness Checkability and Performance Predictability of Systems at Scale
Collaborative Research: PPoSS: LARGE: ScaleStuds: Foundations for Correctness Checkability and Performance Predictability of Systems at Scale
批准号:
2118512
负责人:
Manos Kapritsos
金额:
$62.5万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-10-01 至 2026-09-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
In light of the limits of Moore's Law and Dennard scaling and the ever increasing computing demand, the last decade has seen unprecedented deployment scales; Google is known to run clusters with thousands of machines each, Apple deploys a total of 100,000 database machines, and Netflix runs tens of database clusters with 500 nodes each. This era of extreme-scale distributed systems has given birth to a new class of faults, "scalability faults" -- complex latent faults that are scale-dependent, whose symptoms surface in large-scale deployments but not necessarily in small/medium-scale deployments. Many fundamental research questions are not answerable today. On correctness: How to detect bugs that only manifest under large scale through program analysis? How to test and reproduce various dimensions of system scales efficiently on one machine? How to prevent and fix scalability-related faults? On performance: How to reason about software performance on various heterogeneous devices? How to accurately predict performance of fine-grained tasks to reduce inaccuracies at the aggregate level and project performance to future architectures? Finally, in combination: How to answer all these questions for the larger connected ecosystem -- not just the individual software and hardware components -- and to eventually build future-generation systems that are reproducible and verifiable by construction with respect to correctness and performance at scale? The ScaleStuds project involves a team of ten researchers to develop the foundations of correctness checkability (CC) and performance predictability (PP) of systems at scale. The key principle of this project is to "check large with large" -- check large-scale systems with a large fleet of data, analysis, tests, learning, models, and proofs. The vision is to build an ecosystem of distributed "CC+PP-certified" software-software and -hardware interactions. The project is paving the vision one "floor" at a time, creating composable building blocks ("the studs"). The project first builds new mechanisms such as a scale-testing platform and a unified database of software program properties and hardware performance profiles exposing clear APIs. These studs then enable multi-dimensional automated scalability tests and program analysis and performance learning and prediction at various levels of the software/hardware stack. Ultimately all of these experiences are intended to lead to correct and performant cross-layer/service interactions and future design principles including reproducible- and verified-by-construction development methods. The project novelties include the advancement of debugging, testing, learning, and prediction methods to ensure correctness checkability and performance predictability of extreme-scale systems and applications both on classical hardware platforms and emerging ones; a unified data ecosystem of software/hardware properties and profiles that facilitates automated analyses via clear APIs; a multi-dimensional scale-testing framework that empowers the development of new large-scale unit-tests and program analysis; detailed device profiling and observation to enable large-scale performance learning/prediction and deliver lessons for learning/predicting the behavior of other devices and layers in an end-to-end hardware/software stack; and ultimately a clear definition of CC+PP-certifiability for today's systems and future verifiable/reproducible-by-construction development methods.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1145/3591235
发表时间:
2023
期刊:
Proceedings of the ACM on Programming Languages
影响因子:
--
作者:
[Zhang, Tony Nuda, Sharma, Upamanyu, Kapritsos, Manos]
通讯作者:
Kapritsos, Manos
Collaborative Research: FMitF: Track I: Simplifying End-to-End Verification of High-Performance Distributed Systems
-
批准号:2318954
-
项目类别:Standard Grant
-
资助金额:$37.5万
-
财政年份:2023
-
负责人:Manos Kapritsos
-
依托单位:
CAREER: Formal Verification of Performance Properties for Distributed Systems
-
批准号:2045541
-
项目类别:Continuing Grant
-
资助金额:$56.14万
-
财政年份:2021
-
负责人:Manos Kapritsos
-
依托单位:
FMitF: Track I: Automating the Verification of Distributed Systems
-
批准号:2018915
-
项目类别:Standard Grant
-
资助金额:$74.99万
-
财政年份:2020
-
负责人:Manos Kapritsos
-
依托单位:
CSR: Small: Replication in the Cloud Era
-
批准号:1814507
-
项目类别:Standard Grant
-
资助金额:$49.72万
-
财政年份:2018
-
负责人:Manos Kapritsos
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: