CAREER: Avoiding Achilles' Heel in Exascale Computing with Distributed File Systems
CAREER: Avoiding Achilles' Heel in Exascale Computing with Distributed File Systems
批准号:
1054974
负责人:
Ioan Raicu
金额:
$45.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-01-01 至 2018-06-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Exascale (i.e. 1018 operations/sec) computers will enable the unraveling of significant scientific mysteries, covering many domains (e.g. weather modeling, national security, energy, and drug discovery). Predictions are that exascales will be reached in 2019, with millions of compute-nodes and billions of threads of execution. The current state-of-the-art storage in high-end computing (HEC), in which storage is segregated from compute-nodes and connected by a network (e.g. parallel filesystems), will not scale with the expected exponential growth in concurrency. At exascales, basic functionality (e.g. booting, check-pointing, metadata/data access) at high concurrency levels will suffer poor performance, and combined with system mean-time-to-failure in hours, will lead to a performance collapse. The investigator envisions future HEC systems to be designed with non-volatile memory on every compute node, and every node to actively participate in the metadata and data management. This work aims to: 1) design, analyze, and implement a distributed data structure (D3) optimized for HEC, to be used for distributed metadata management; 2) design, analyze, and implement a distributed filesystem (FDFS) optimized for a subset of important high-performance computing (HPC) as well as many-task computing (MTC) workloads, and scalable to millions of nodes; and 3) evaluate work with real workloads, applications, and simulations up to exascales. The results of this work has the potential to make exascale computing more tractable, touching virtually all disciplines in HEC, fueling scientific discovery and economic development at the national level. The HEC knowledgebase will extend into commodity systems as the fastest machines generally become mainstream systems in five to seven years. This work can also open doors for research in radical parallel programming paradigms (e.g. MTC) that rely on scalable storage infrastructure.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: REU Site: BigDataX: From theory to practice in Big Data computing at eXtreme scales
-
批准号:2150500
-
项目类别:Standard Grant
-
资助金额:$36.29万
-
财政年份:2022
-
负责人:Ioan Raicu
-
依托单位:
Collaborative Research: OAC Core: Enabling Extremely Fine-grained Parallelism on Modern Many-core Architectures
-
批准号:2107548
-
项目类别:Standard Grant
-
资助金额:$33.37万
-
财政年份:2021
-
负责人:Ioan Raicu
-
依托单位:
REU Site: Collaborative Research: BigDataX: From theory to practice in Big Data computing at eXtreme scales
-
批准号:1757964
-
项目类别:Standard Grant
-
资助金额:$32.31万
-
财政年份:2018
-
负责人:Ioan Raicu
-
依托单位:
CRI: II-NEW: MYSTIC: Programmable Systems Research Testbed to Explore a Stack-WIde Adaptive System fabriC
-
批准号:1730689
-
项目类别:Standard Grant
-
资助金额:$100.0万
-
财政年份:2017
-
负责人:Ioan Raicu
-
依托单位:
REU Site: BigDataX: From Theory to Practice in Big Data Computing at Extreme Scales
-
批准号:1461260
-
项目类别:Standard Grant
-
资助金额:$28.8万
-
财政年份:2015
-
负责人:Ioan Raicu
-
依托单位:
Student Travel Support for ACM HPDC 2011
-
批准号:1114379
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2011
-
负责人:Ioan Raicu
-
依托单位:
海外基金