课题基金 / 基金详情

CAREER: Rethinking HPC Resilience in the Exascale Era

CAREER: Rethinking HPC Resilience in the Exascale Era
职业:重新思考百亿亿次时代的 HPC 弹性
批准号:
1750503
负责人:
Changhee Jung
金额:
$52.17万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-01-15 至 2019-11-30

项目摘要

项目成果

Changhee Jung的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Resilience is one of the key exascale research challenges in high-performancecomputing (HPC). Due to much high error rates, exascale supercomputers couldmake little progress in computations, or might generate incorrect results due tofailures, rendering the exascale performance useless. Thechallenge is how to achieve a complete HPC resilience at exascale in a way thatdoes not increase the performance overhead, the power consumption, and thecomplexity of underlying hardware. To this end, this research project designsand develops low-cost hardware/software cooperative techniques for HPCresilience in the exascale era. This project involves four research goals: (1) low-cost soft error resiliencefor CPUs; intelligent compiler-architecture interaction can validate the lack oferrors and performs fine-grained recovery, thus eliminating SDC. (2)compiler-directed soft error resilience for commodity GPUs; it can remove thepower-hungry error-correcting code (ECC) logic from the GPU register fileswithout compromising their resilience. (3) lightweight nonvolatile memory (NVM)persistence; it can mitigate the overhead of traditional heavyweight HPCcheckpointing and support whole-system persistence for applications withoutirrevocable operations. (4) low-cost timing error resilience for aggressivevoltage scaling to maximize the energy-efficiency with program correctnessguarantee.The resulting artifacts and technologies are expected to contribute to thenation's competitiveness by addressing the challenge of building reliable HPCsystems. The research outcome impacts a broad range of any disciplines thatneed correct computation results thus requiring reliable computing systemscovering from embedded systems to HPC cloud. Consequently, use of the proposedtechniques will make the execution of current and emerging applications muchmore reliable, and therefore directly affect our way of life.There will be three types of data generated from this research project: (1)algorithms and models, (2) software prototype, (3) testing infrastructureincluding simulators and evaluation benchmarks and their traces, (4) educationalmaterials. All of our software tools will be open source and made available tothe public, laboratories and industry.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
CommAnalyzer: Automated Estimation of Communication Cost and Scalability on HPC Clusters from Sequential Code
CommAnalyzer:根据顺序代码自动估计 HPC 集群的通信成本和可扩展性
DOI: --
发表时间: 2018
期刊: ACM International Symposium on High-Performance Parallel and Distributed Computing (HPDC
影响因子: --
作者: [Helal, Ahmed, Jung, Changhee, Feng, Wu-chun, Hanafy, Yasser]
通讯作者: Hanafy, Yasser
Collaborative Research: CSR: Small: Caphammer: A New Security Exploit in Energy Harvesting Systems and its Countermeasures
  • 批准号:
    2314681
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $24.0万
  • 财政年份:
    2023
  • 负责人:
    Changhee Jung
  • 依托单位:
Collaborative Research: SHF: Small: Enabling Caches and GPUs for Energy Harvesting Systems
  • 批准号:
    2153749
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2022
  • 负责人:
    Changhee Jung
  • 依托单位:
CAREER: Rethinking HPC Resilience in the Exascale Era
  • 批准号:
    2001124
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.67万
  • 财政年份:
    2019
  • 负责人:
    Changhee Jung
  • 依托单位:
SHF: Small: Compiler and Architectural Techniques for Soft Error Resilience
海外基金