课题基金 / 基金详情

ITR: Scalable Non-Stop Blade-Based Servers

ITR: Scalable Non-Stop Blade-Based Servers
ITR:可扩展的不间断刀片服务器
批准号:
0325802
负责人:
James Hoe
金额:
$120.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-10-01 至 2007-09-30

项目摘要

项目成果

James Hoe的其他基金

相似基金

相关文献

中文摘要
翻译
该项目提出了可扩展的Non-stOP服务器(SNOPS)架构:一个可靠、可用和可服务(RAS)的硬件平台。SNOPS通过使用通过可扩展网络和硬件分布式共享内存(DSM)互连的商用刀片组件,提供了传统面向RAS的服务器所无法比拟的成本和性能可扩展性。SNOPS通过对应用程序透明的快速检测和恢复软/瞬时错误和/或单个处理器/存储器组件故障,以及在硬件组件故障时热插拔模块而不丢失状态,提供不间断的服务。该项目提出了跨网络的存储器总线(MEMBRANE),一个抽象层来调节处理器和存储器模块之间的存储器数据传输。MEMBRANE是处理器和存储器端错误数据的不可穿透的屏障,只允许无错误数据传输。在处理器侧,MEMBRANE一致性协议通过比较它们的请求来检测和触发对源自一组冗余处理器的错误的恢复。在内存方面,MEMBRANE内存冗余协议通过类似RAID的分布式奇偶校验方案检测并恢复来自内存组件的错误。该原型将支持商业操作系统(如Linux),并提供必要的性能,以允许针对商业级服务器应用程序对我们的想法进行全面评估。可扩展的服务器模拟基础设施也将可用于分发,允许快速,准确和全系统的服务器模拟。在工业和学术环境中已经建立了几个概念验证原型,特别是对学生来说,这是一个宝贵的经验,但也有很大的潜力对工业研究和开发产生影响。
英文摘要
This project proposes the Scalable Non-stOP Server (SNOPS) architecture: a reliable, available, and serviceable (RAS) hardware platform. SNOPS offers both cost and performance scalability unparalleled by conventional RAS-oriented servers by using commodity blade components interconnected through a scalable network and hardware distributed shared memory (DSM). SNOPS offers non-stop service through fast application-transparent detection and recovery of soft/transient errors and/or single processor/memory component failure, and hot-swapping of a module upon hardware component failure without state loss.The project proposes the MEMory BaRrier Across the NEtwork (MEMBRANE), an abstraction layer to regulate memory data transfer between processor and memory modules. MEMBRANE serves as an impenetrable barrier for faulty data both from the processor and the memory sides, allowing only error-free data to transfer across. On the processor side, the MEMBRANE coherence protocols detect and trigger recovery for errors originating from redundant processors of a group by comparing their requests. On the memory side, the MEMBRANE memory redundancy protocols detect and recover from errors originating from the memory components via a RAID-like distributed parity scheme.This prototype will support a commodity OS (such as Linux) and deliver the necessary performance to permit full-scale evaluation of our ideas against commercial-grade server applications. Scalable server simulation infrastructure will be also be available for distribution, allowing for fast, accurate, and full-system simulation of servers. Several proof-of-concept prototypes have been built in industrial and academic settings and especially for students are an invaluable experience, but with a high potential to have impact on industrial research and development as well.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CCF: Small: Accelerating Irregular Algorithms using Cache-Coherent FPGA Accelerators
  • 批准号:
    1618014
  • 项目类别:
    Standard Grant
  • 资助金额:
    $33.0万
  • 财政年份:
    2016
  • 负责人:
    James Hoe
  • 依托单位:
SHF: Small: Compiling Custom Hardware Accelerators from Graph Algorithms
  • 批准号:
    1320725
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2013
  • 负责人:
    James Hoe
  • 依托单位:
SHF: Large: Rethinking the Architecture of FPGAs as First-Class Computing Devices
  • 批准号:
    1012851
  • 项目类别:
    Standard Grant
  • 资助金额:
    $100.0万
  • 财政年份:
    2010
  • 负责人:
    James Hoe
  • 依托单位:
CPA-CSA: Accelerating Architectural-level, Full-system Multiprocessor Simulations using FPGAs
  • 批准号:
    0811702
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2008
  • 负责人:
    James Hoe
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis