课题基金 / 基金详情

MRI: Acquisition of a Heterogeneous Multi-GPU Cluster to Support Exploration at Scale

MRI: Acquisition of a Heterogeneous Multi-GPU Cluster to Support Exploration at Scale
MRI:获取异构多 GPU 集群以支持大规模探索
批准号:
1920020
负责人:
David Kaeli
金额:
$39.97万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2021-09-30

项目摘要

项目成果

David Kaeli的其他基金

相似基金

相关文献

中文摘要
翻译
该项目旨在获得一个异构的多GPU集群,由最先进的GPU设备构建,与新兴的NVLink和HDR网络互连,用于GPU缓存的网络附加非易失性存储器(NVM)存储,并通过智能HDR Infiniband交换机互连,以启用,加速,探索和支持来自不同领域的大规模应用程序,包括:·分布式视网膜病变深度神经网络,·无线网络取证,·对抗机器学习,·计算社会科学,·数学优化和大数据分析,·海岸工程建模,·多GPU系统(包括支持GPU网络缓存的NVMe技术和可以卸载集体操作的智能网络交换机)这些功能将使计算科学家能够通过编程智能网络交换机和选择性缓存来隐藏内存和互连延迟,以新的方式利用GPU并行性。目前,图形处理单元(GPU)通过重叠计算和存储器操作来启动大量线程,从而提供高计算吞吐量。结合低开销的线程交换,GPU可以隐藏长内存操作。但是,底层系统架构并没有跟上GPU应用程序的规模和复杂性的增长。与多CPU系统相比,多GPU解决方案对程序员不太友好,并且在其架构支持方面导致可扩展性较低。当前的GPU系统将GPU视为分立设备,对真正共享内存编程模型的支持有限。由于多GPU互连带宽已经成为扩展多GPU系统、探索新网络拓扑、更智能的网络元件以及用于缓存和预取的增强软件层的限制因素,该奖项反映了NSF的法定使命,并通过利用基金会的知识价值和更广泛的影响进行评估,被认为值得支持审查标准。
英文摘要
This project aims to acquire a heterogeneous Multi-GPU cluster, constructed out of state-of-the-art GPUs devices, interconnected with emerging NVLink and HDR networks, network-attached non-volatile memory (NVM) storage for GPU caching, and interconnected by a smart HDR infiniband switch, to enable, accelerate, explore, and support applications at scale from different domains that include:• Distributed deep neural networks for retinopathy,• Wireless network forensics,• Adversarial machine learning,• Computational social science,• Mathematical optimization and big data analytics,• Coastal engineering modeling, and• Multi-GPU system (including NVMe technology to support caching in GPU network and a smart network switch that can offload collective operations)These features will enable computational scientists to exploit GPU parallelism in new ways by programming the smart network switch and caching selectively to hide memory and interconnect latency.Currently, graphics processing units (GPUs) provide high computational throughput by lunching a large number of threads by overlapping compute and memory operations. Combined with low-overhead thread swapping, GPUs can hide long memory operations. But underlying system architectures have not kept up as the size and complexity of GPU applications grow. The multi-GPU solutions are less programmer friendly and result in lower scalability when their architectural support is compared with the multi-CPU systems. Current GPUs systems treat GPUs as discrete devices, with limited support for a truly shared memory programming model. Since multi-GPU interconnect bandwidth has become a limiting factor for scaling multi-GPU systems, exploration of new network topologies, smarter network elements, and enhanced software layers for caching and prefetching, that meet the needs of tomorrow’s demanding data applications are necessary.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CSR: Medium: Architecting GPUs for Practical Homomorphic Encryption-based Computing
  • 批准号:
    2312275
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $60.0万
  • 财政年份:
    2023
  • 负责人:
    David Kaeli
  • 依托单位:
REU Site: REU Research Experiences and Mentoring in Data-Driven Discovery
  • 批准号:
    1559894
  • 项目类别:
    Standard Grant
  • 资助金额:
    $35.96万
  • 财政年份:
    2016
  • 负责人:
    David Kaeli
  • 依托单位:
Student Travel for PACT 2016
  • 批准号:
    1624175
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2016
  • 负责人:
    David Kaeli
  • 依托单位:
STARSS: Small: Side-Channel Analysis and Resiliency Targeting Accelerators
  • 批准号:
    1618379
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2016
  • 负责人:
    David Kaeli
  • 依托单位:
海外基金