课题基金 / 基金详情

Collaborative Research: FuSe: R3AP: Retunable, Reconfigurable, Racetrack-Memory Acceleration Platform

Collaborative Research: FuSe: R3AP: Retunable, Reconfigurable, Racetrack-Memory Acceleration Platform
合作研究:FuSe:R3AP:可重调、可重新配置、赛道内存加速平台
批准号:
2328974
负责人:
Kang Wang
金额:
$94.5万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-01-01 至 2026-12-31

项目摘要

项目成果

Kang Wang的其他基金

相似基金

相关文献

中文摘要
翻译
在传统的冯·诺依曼计算系统中,由于计算单元之间的数据传输速度大大落后于容量、处理速度和效率,因此出现了一个显著的瓶颈。为了通过弥合存储和计算之间的差距来缓解这一瓶颈,已经引入了许多创新的存储技术,以及为新兴和传统内存系统设计的近内存和内存处理解决方案。然而,一个相当大的挑战仍然存在:实际制造系统的原型和特征,特别是那些包含成熟技术和尖端技术的系统。为了克服这一挑战,该项目开发了一种基于新兴赛道内存的尖端可返回和可重构加速平台(R3AP),利用了设备-架构-应用协同设计方法。R3AP的突出特性包括其作为可重构逻辑、内存中处理(PIM)加速器和高密度内存存储的功能。它是可返回的,这意味着它可以使用位、整数和浮点精度进行操作,并且可以模拟类似模拟的存储和处理。R3AP有效地缓解了数据移动的低效率,同时提供特定于领域的加速和适应性。凭借其密集、可靠、节能和超低延迟的计算能力,R3AP有可能彻底改变未来计算系统的存储和处理能力,例如物联网(IoT)和网络物理系统(CPS)。它还可以应用于高性能和云计算系统。该项目的研究成果通过出版物、研讨会、设计竞赛、教程、工业课程和技术转让活动进行分享。教育资源和推广活动计划在项目网站上提供,软件工件在GitHub上发布。为了实现R3AP,该项目包括一系列相互关联的研究任务,跨越多个系统层。在器件层面,该项目将电压控制的skyrmion运动机制与工业级8英寸晶圆磁性隧道结堆栈集成在一起,并展示了全功能的skyrmion赛道存储器(SRTM),包括skyrmion流的形成、移动和检测。此外,本文还评估了SRTM的性能,重点关注写错误率、移错误率、读错误率、操作速度和能耗等方面。它还解决并减轻了非理想性,例如钉住效应,并继续开发和演示cmos集成SRTM。在体系结构和电路层,该项目涉及创建可变查找表、计算和存储单元。该单元的性能类似于多上下文现场可编程门阵列(FPGA)逻辑、并行PIM逻辑、大规模并行累加器以及类似模拟的存储和计算结构,利用了SRTM的独特属性。这一层确保了来自由银行、子阵列、块等组成的层次结构的高速内存访问,并进一步通过可配置的开关箱和基于网格的片上网络添加链接,以支持PIM的数据移动操作,否则这将是具有挑战性的。在应用层,该项目开发了新颖的建模、分析、设计空间探索和运行时调整技术,以利用R3AP提供的高度可重构性。目标是使未来的物联网和CPS应用适应不断变化的环境和需求,优化资源使用,抵御外部干扰,提高整体系统性能、弹性和可持续性。在所有这些层中,该项目开发了一个可扩展的计算机辅助设计(CAD)流程。这涉及到一个基于多级中间表示的编译流,它可以将高级描述语言(如PyTorch和C/ c++)编译为R3AP设备的二进制文件。该流程使用多级层次结构,包括设计的前端、中端和后端编译,并将各种优化和管理问题抽象到合适的级别,以便有效解决。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In traditional Von Neumann computing systems, a significant bottleneck arises because the data transfer speed to and from the computing units has considerably fallen behind capacity, processing speed, and efficiency. To mitigate this bottleneck by bridging the gap between storage and computation, many innovative storage technologies have been introduced, along with near- and in-memory processing solutions designed for both emerging and traditional memory systems. Nonetheless, a considerable challenge remains: the prototyping and characterization of actual fabricated systems, especially those encompassing both mature technologies and cutting-edge technologies. To overcome this challenge, this project develops a cutting-edge Retunable and Reconfigurable Acceleration Platform (R3AP) based on emerging racetrack memory, leveraging a device-architecture-application co-design approach. The standout features of R3AP include its ability to function as a reconfigurable logic, a processing-in-memory (PIM) accelerator, and a high-density memory storage. It is retunable, meaning it can operate with bit-wise, integer, and floating-point precision, and can simulate analog-like storage and processing. R3AP effectively mitigates data movement inefficiencies while offering domain-specific acceleration and adaptability. With its dense, reliable, energy-efficient, and ultra-low latency computational capability, R3AP has the potential to revolutionize the storage and processing capabilities of future computing systems, such as those in Internet of Things (IoT) and Cyber-Physical Systems (CPS). It can also be applied to high-performance and cloud computing systems. The project's findings are shared through publications, workshops, design contests, tutorials, industrial courses, and technology transfer activities. Educational resources and outreach activity plans are made available on the project website, and software artifacts are released on GitHub.To realize R3AP, the project comprises a series of interrelated research tasks spanning multiple system layers. At the device level, the project integrates the voltage-controlled skyrmion motion mechanism with the industrial-grade 8-inch wafer magnetic tunneling junction stack and demonstrates a fully functional Skyrmion racetrack memory (SRTM), including the formation, shifting, and detection of the skyrmion stream. Additionally, it evaluates the performance of SRTM, focusing on aspects such as write-error-rate, shift-error-rate, read-error-rate, operation speed, and energy consumption. It also addresses and mitigates non-idealities, such as the pinning effect, and goes on to develop and demonstrate CMOS-integrated SRTM. On the architecture and circuit layers, the project involves the creation of a mutable lookup table, compute, and memory unit. This unit performs like multi-context Field-Programmable Gate Array (FPGA) logic, parallel PIM logic, massively parallel accumulators, and analog-like storage and compute structures, leveraging the unique properties of SRTM. This layer ensures high-speed memory access from a hierarchy consisting of banks, subarrays, tiles, etc., and further adds links via configurable switch boxes and a mesh-based network-on-chip to enable data movement operations for PIM that would otherwise be challenging. At the application layer, the project develops novel modeling, analysis, design space exploration, and runtime adjustment techniques to exploit the high degree of reconfigurability provided by R3AP. The goal is to adapt future IoT and CPS applications to changing environments and requirements, optimize resource usage, withstand external disturbances, and enhance overall system performance, resilience, and sustainability. Across all these layers, the project develops a scalable computer-aided design (CAD) flow. This involves a multi-level intermediate representation-based compilation flow, which can compile high-level description languages such as PyTorch and C/C++ into binaries for the R3AP device. This flow uses a multi-level hierarchy including front-end, middle-end, and back-end compilation of the designs, and abstracts various optimization and management problems to a suitable level for efficient resolution.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Workshop on Future of Semiconductors and Beyond: Devices and Technologies: To Be Held Virtually Feb 8-9, and 17-18, 2021.
NSF Convergence Accelerator Track C: Chiral-Based Quantum Interconnect Technologies (CirquiTs)
SHF: Small: Collaborative Research: Skyrmion Mediated Energy-efficient VCMA Switching of 2-Terminal p-MTJ Memory
Ultra-fast energy efficient ferrimagnetic memory
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)