课题基金 / 基金详情

HECURA: Toward Automated Problem Analysis of Large Scale Storage Systems

HECURA: Toward Automated Problem Analysis of Large Scale Storage Systems
HECURA:迈向大规模存储系统的自动化问题分析
批准号:
0621508
负责人:
Priya Narasimhan
金额:
$99.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-07-15 至 2009-06-30

项目摘要

项目成果

Priya Narasimhan的其他基金

相似基金

相关文献

中文摘要
翻译
CMU建议探索用于自动分析大规模存储系统中的故障和性能降级的方法和算法。问题分析包括关键任务,如确定哪个组件(S)行为不当,可能的根本原因,以及支持任何结论的证据。通过将统计工具与适当的检测手段相结合,我们希望大幅降低分析已部署存储系统中的性能和可靠性问题的难度。这些工具与自动化反应逻辑相结合,也为自我修复的长期目标提供了重要的构建块。自动化问题分析对于实现未来高端计算系统所需的规模的成本效益存储至关重要。硬件和软件组件的数量将使问题变得常见而不是异常,因此必须能够快速地从问题转移到修复,而几乎不需要系统停机来进行分析。此外,这种系统的分布式软件复杂性使手工分析变得越来越站不住脚。更微妙但可能是最令人担忧的是,这些存储系统的实施者越来越无法在典型的高端计算环境中进行测试,因为他们根本负担不起重建必要的系统规模。因此,必须在外地分析与规模有关的问题,以便进行改进,这会给客户/用户带来延误和生产力降低,以及为支持高度敏感的活动而部署的系统的许可问题。目前的设计和工具远远达不到所需。
英文摘要
CMU proposes to explore methodologies and algorithms for automating analysis of failures and performance degradations in large-scale storage systems. Problem analysis includes such crucial tasks as identifying which component(s) misbehaved, likely root causes, and supporting evidence for any conclusions. Combining statistical tools with appropriate instrumentation, we hope to dramatically reduce the difficulty of analyzing performance and reliability problems in deployed storage systems. Such tools, integrated with automated reaction logic, also provide an essential building block for the longer-term goal of self-healing.Automating problem analysis is crucial to achieving cost-effective storage at the scales needed for tomorrow's high-end computing systems. The number of hardware and software components will make problems common rather than anomalous, so it must be possible to quickly move from problem to fix with little-to-no system downtime for analysis. Further, the distributed software complexity of such systems make by-hand analysis increasingly untenable. More nuanced, but perhaps of most concern, implementors of these storage systems are increasingly unable to test in representative high-end computing environments because they simply cannot afford to recreate the necessary system scale. As a result, scale-related problems must be analyzed in the field to allow improvements to be made, which introduces delays and productivity reductions for customers/users plus issues of clearance for systems deployed to support highly sensitive activities. Current designs and tools fall far short of what is needed.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Integrated Fault Tolerance and Real-Time Support for Middleware Applications
  • 批准号:
    0238381
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2003
  • 负责人:
    Priya Narasimhan
  • 依托单位:
国内基金
海外基金
Toward a general theory of intermittent aeolian and fluvial nonsuspended sediment transport
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    55万元
  • 批准年份:
    2022
  • 负责人:
    Thomas Pahtz
  • 依托单位: