课题基金 / 基金详情

CAREER: Avoiding Achilles' Heel in Exascale Computing with Distributed File Systems

CAREER: Avoiding Achilles' Heel in Exascale Computing with Distributed File Systems
职业:使用分布式文件系统避免百亿亿次计算中的致命弱点
批准号:
1054974
负责人:
Ioan Raicu
金额:
$45.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-01-01 至 2018-06-30

项目摘要

项目成果

Ioan Raicu的其他基金

相似基金

相关文献

中文摘要
翻译
亿级(即1018次/秒)计算机将能够揭开重大科学谜团,涵盖许多领域(例如天气建模、国家安全、能源和药物发现)。据预测,2019年将达到亿级规模,拥有数百万个计算节点和数十亿个执行线程。目前最先进的高端计算存储(HEC)将存储与计算节点隔离,并通过网络(例如并行文件系统)连接,不会随着并发性的预期指数增长而扩展。在亿级级别上,高并发级别的基本功能(如引导、检查点、元数据/数据访问)将遭受较差的性能,并且与以小时为单位的系统平均故障时间相结合,将导致性能崩溃。研究人员设想,未来的HEC系统将在每个计算节点上设计非易失性存储器,并在每个节点上积极参与元数据和数据管理。这项工作的目标是:1)设计、分析和实现一个针对HEC优化的分布式数据结构(D3),用于分布式元数据管理;2)设计、分析和实现一个分布式文件系统(FDFS),该系统针对重要的高性能计算(HPC)和多任务计算(MTC)工作负载的子集进行优化,并且可扩展到数百万个节点;以及3)评估具有实际工作负载、应用程序和模拟的工作,最高可达数千级。这项工作的结果有可能使亿级计算变得更容易处理,几乎涉及到HEC的所有学科,推动国家一级的科学发现和经济发展。HEC的知识库将扩展到商品系统,因为最快的机器通常在五到七年内成为主流系统。这项工作还可以为依赖于可扩展存储基础设施的激进并行编程范例(例如MTC)的研究打开大门。
英文摘要
Exascale (i.e. 1018 operations/sec) computers will enable the unraveling of significant scientific mysteries, covering many domains (e.g. weather modeling, national security, energy, and drug discovery). Predictions are that exascales will be reached in 2019, with millions of compute-nodes and billions of threads of execution. The current state-of-the-art storage in high-end computing (HEC), in which storage is segregated from compute-nodes and connected by a network (e.g. parallel filesystems), will not scale with the expected exponential growth in concurrency. At exascales, basic functionality (e.g. booting, check-pointing, metadata/data access) at high concurrency levels will suffer poor performance, and combined with system mean-time-to-failure in hours, will lead to a performance collapse. The investigator envisions future HEC systems to be designed with non-volatile memory on every compute node, and every node to actively participate in the metadata and data management. This work aims to: 1) design, analyze, and implement a distributed data structure (D3) optimized for HEC, to be used for distributed metadata management; 2) design, analyze, and implement a distributed filesystem (FDFS) optimized for a subset of important high-performance computing (HPC) as well as many-task computing (MTC) workloads, and scalable to millions of nodes; and 3) evaluate work with real workloads, applications, and simulations up to exascales. The results of this work has the potential to make exascale computing more tractable, touching virtually all disciplines in HEC, fueling scientific discovery and economic development at the national level. The HEC knowledgebase will extend into commodity systems as the fastest machines generally become mainstream systems in five to seven years. This work can also open doors for research in radical parallel programming paradigms (e.g. MTC) that rely on scalable storage infrastructure.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: REU Site: BigDataX: From theory to practice in Big Data computing at eXtreme scales
  • 批准号:
    2150500
  • 项目类别:
    Standard Grant
  • 资助金额:
    $36.29万
  • 财政年份:
    2022
  • 负责人:
    Ioan Raicu
  • 依托单位:
Collaborative Research: OAC Core: Enabling Extremely Fine-grained Parallelism on Modern Many-core Architectures
  • 批准号:
    2107548
  • 项目类别:
    Standard Grant
  • 资助金额:
    $33.37万
  • 财政年份:
    2021
  • 负责人:
    Ioan Raicu
  • 依托单位:
REU Site: Collaborative Research: BigDataX: From theory to practice in Big Data computing at eXtreme scales
  • 批准号:
    1757964
  • 项目类别:
    Standard Grant
  • 资助金额:
    $32.31万
  • 财政年份:
    2018
  • 负责人:
    Ioan Raicu
  • 依托单位:
CRI: II-NEW: MYSTIC: Programmable Systems Research Testbed to Explore a Stack-WIde Adaptive System fabriC
  • 批准号:
    1730689
  • 项目类别:
    Standard Grant
  • 资助金额:
    $100.0万
  • 财政年份:
    2017
  • 负责人:
    Ioan Raicu
  • 依托单位:
海外基金