课题基金 / 基金详情

CSR: EAGER: An Integrated Framework for Performance and Reliability in Large-scaled Computing Systems

CSR: EAGER: An Integrated Framework for Performance and Reliability in Large-scaled Computing Systems
CSR:EAGER:大规模计算系统性能和可靠性的集成框架
批准号:
1251129
负责人:
Ningfang Mi
金额:
$27.24万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-09-01 至 2014-08-31

项目摘要

项目成果

Ningfang Mi的其他基金

相似基金

相关文献

中文摘要
翻译
数据中心和云计算等大规模计算环境正在成为核心计算基础设施,这使得此类服务的可用性变得极其关键。然而,这些环境越来越容易受到硬件和软件故障的影响。该项目设计了故障感知技术,用于在存在各种级别的硬件和软件故障的大规模计算环境中进行建模、预测和资源管理。在智力上,这个项目发展了对工作量和可靠性特征的基本理解,并调查了改进的容量规划模型和预测技术如何能够为系统设计和维护获得有用的信息。该项目进一步深入了解了软件/硬件组件故障对资源管理领域的影响。该项目的结果将包括评估给定系统的可靠性和性能的新容量规划模型,以及通过利用故障事件中的时间相关性来预测未来故障发生的新预测技术。基于建模和预测技术,该项目将开发新的故障感知运行时策略,用于作业调度、节点分配和系统维护,旨在实现复杂大型系统的高性能和可靠性。
英文摘要
Large-scale computing environments such as data centers and cloud computing are becoming the core computing infrastructure, making the availability of such services extremely critical. However, these environments are increasingly vulnerable to both hardware and software failures. This project designs failure-aware techniques for modeling, prediction, and resource management in large-scale computing environments with the presence of hardware and software failures at various levels. Intellectually, this project develops fundamental understanding of workload and reliability characteristics, and investigates how improved capacity planning models and prediction techniques can obtain useful information for system design and maintenance. This project further provides insights of the impact of software/hardware component failures in the area of resource management. The results of this project will include new capacity planning models that evaluate both reliability and performance of a given system and new prediction techniques that forecast the future failure occurrences by taking advantage of temporal dependence in failure events. Based on the modeling and prediction techniques, this project will develop new failure-aware runtime strategies for job scheduling, node allocation, and system maintenance, aiming to achieve high system performance and reliability in complex large scale systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CNS core: OAC core: Small: New Techniques for I/O Behavior Modeling and Persistent Storage Device Configuration
  • 批准号:
    2008072
  • 项目类别:
    Standard Grant
  • 资助金额:
    $24.49万
  • 财政年份:
    2020
  • 负责人:
    Ningfang Mi
  • 依托单位:
CAREER: Capacity Planning Methodologies for Large Clusters with Heterogeneous Architectures and Diverse Applications
  • 批准号:
    1452751
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $45.96万
  • 财政年份:
    2015
  • 负责人:
    Ningfang Mi
  • 依托单位:
海外基金