CAREER: Storage-Aware Fault Tolerance
CAREER: Storage-Aware Fault Tolerance
批准号:
2339784
负责人:
Aishwarya Ganesan
金额:
$69.96万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-02-15 至 2029-01-31
中文摘要
容错存储系统是现代数据中心的核心。这些系统确保大型服务(例如,搜索、社交网络)和关键服务(例如,电子商务、医疗保健)即使在出现故障时也可以可靠地访问基本数据。然而,在设计容错存储系统时,一个紧迫的问题是它们强大的容错保证是以性能为代价的。例如,使用现有方法构建的容错键值存储的性能可能比存储的独立版本(不能容忍故障)低20倍。该项目旨在构建接近容错(非复制)存储服务器性能的容错存储系统。该项目将通过系统地重新思考广泛使用的容错范例来实现这一目标。该项目将开发新的容错协议和抽象,并建立新的实用系统。特别是,它将首先开发一种针对现代存储设备进行优化的新复制协议。其次,它将探索一种新的无CPU复制方法,以充分释放远程直接内存访问(RDMA)的潜力。第三,它将实现一个新的容错架构,专为新兴的分散式数据中心量身定做。最后,它将为存储应用程序开发一种新的共享日志抽象。在该项目中开发的解决方案将使开发可靠和高性能的系统成为可能,消除了威胁关键应用程序数据安全的妥协需要。这一努力将通过新课程向学生介绍分布式系统研究和动手实验室,为学生使用现代硬件做好准备,从而极大地促进教育和推广。该项目还将通过博士研讨会和参与本科生研究来扩大对计算机的参与。该项目还将通过新的交互框架将分布式系统带给更广泛的受众(包括K-12学生)。为了方便使用,项目中的所有构件都将公开提供,并提供必要的文档。最后,PI将与行业合作伙伴合作,在现实世界的系统中实施项目成果。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Fault-tolerant storage systems are at the heart of modern datacenters. These systems ensure that large-scale services (e.g., search, social networking) and critical services (e.g., e-commerce, healthcare) can reliably access essential data, even in the face of failures. However, a pressing concern when designing fault-tolerant storage systems is that their strong fault-tolerance guarantees come at the cost of performance. For instance, a fault-tolerant key-value store built using existing approaches can perform up to 20x worse than a stand-alone version of the store (that does not tolerate failures). This project aims to build fault-tolerant storage systems that closely approximate the performance of fault-intolerant (non-replicated) storage servers. The project will achieve this goal by systematically rethinking widely used fault-tolerance paradigms. This project will develop novel fault-tolerance protocols and abstractions and build new practical systems. In particular, it will first develop a new replication protocol optimized for modern storage devices. Second, it will explore a novel CPU-free replication approach that unlocks the full potential of remote direct memory access (RDMA). Third, it will realize a new fault-tolerance architecture tailored for emerging disaggregated datacenters. Finally, it will develop a new shared log abstraction for storage applications. The solutions developed in this project will enable the development of reliable and performant systems, eliminating the need to make compromises that threaten the data safety of critical applications. The effort will significantly contribute to education and outreach through new course offerings to introduce students to distributed systems research and hands-on labs to equip students to use modern hardware. The project will also broaden participation in computing through doctoral workshops and engagement in undergraduate research. The project will also bring distributed systems to a broader audience (including K-12 students) through new interactive frameworks. All the artifacts from the project will be made openly available with necessary documentation for ease of use. Finally, the PI will collaborate with industry partners to implement project results within real-world systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
面向 In-Storage 智能计算的高性能 SSD 控制器研究
-
批准号:ZCLJHSQY26F0401
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:何越
-
依托单位:
面向in-storage智能计算的固态硬盘缓存管理优化
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2022
-
负责人:廖剑伟
-
依托单位: