CAREER: Capacity Planning Methodologies for Large Clusters with Heterogeneous Architectures and Diverse Applications
CAREER: Capacity Planning Methodologies for Large Clusters with Heterogeneous Architectures and Diverse Applications
批准号:
1452751
负责人:
Ningfang Mi
金额:
$45.96万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-04-01 至 2022-03-31
中文摘要
该项目侧重于开发创新技术和算法,以建立适当的系统模型,并支持性能和可靠性分析,以便更好地解释大型系统行为,预测应用程序性能,并确保高资源效率和系统可靠性。大型集群环境是当今计算基础设施的重要组成部分,它为运行处理核心业务和操作数据的应用程序提供了平台。然而,随着计算和应用程序基础设施复杂性的增加以及对高质量服务的需求的增长,大型集群环境面临着确保应用程序始终可用并提供足够性能的艰巨任务。该项目期望为大型集群系统的性能建模、工作负载测量和模型参数化实现新的容量规划技术。智能容量和可靠性建模将使服务提供商能够在部署和运行应用程序之前确定其应用程序的最佳平台。它还将使系统管理人员能够优化整个集群基础设施的性能、可靠性和效率。本研究将开发新的性能建模方法,以捕获异构硬件架构的特征,并预测在一系列计算平台上运行的应用程序的行为。该研究将把性能建模扩展到故障感知。改进的模型将通过捕获系统工作负载和故障事件的特征,实现对复杂大型系统的性能和可靠性的准确预测。此外,研究人员将开发新的先进技术,将计算和通信组件的基本处理信息参数化性能模型。这些基本的处理信息不仅限于平均值,而且还包括其他关键但复杂的特征,如资源争用和突发症状。该项目涉及从中学到研究生院的学生的教育活动,以积极激励学生,特别是妇女,将科学和工程研究纳入课程开发和本科生研究活动。
英文摘要
This project focuses on developing innovative techniques and algorithms to build adequate system models and support performance and reliability analysis in order to better explain large system behavior, predict application performance, and ensure high resource efficiency and system dependability. Large cluster environments are an important part of today's computing infrastructure, providing the platform for running applications that handle core business and operational data. However, with the complexity of computing and application infrastructure increasing and the requirements for high quality of service growing, large cluster environments are facing the difficult task of ensuring that applications are always available and delivering adequate performance. This project expects to achieve new capacity planning techniques for performance modeling, workload measurements and model parameterizations of large cluster systems. Intelligent capacity and reliability modeling will enable service providers to determine the best platform for their application before deploying and running the application. It will also enable system managers to optimize the performance, reliability and efficiency of the entire cluster infrastructure. This research will develop new performance modeling methods to capture the characteristics of heterogeneous hardware architectures and predict the behavior of an application running on an array of computing platforms. The research will extend performance modeling to failure awareness. The improved models will enable an accurate prediction of performance and reliability of a complex large-scale system by capturing the characteristics of both system workloads and failure events. In addition, the researchers will develop new advanced techniques to parameterize performance models with essential processing information of computational and communication components. These essential processing information do not only limit to mean values but also include other critical yet complicated features such as resource contention and burstiness symptoms. The project is involved with educational activities reaching out to students from secondary to graduate schools to aggressively motivate students, especially women, towards science and engineering integrating this research into curriculum development and undergraduate research activities.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1109/ipccc50635.2020.9391566
发表时间:
2020-11
期刊:
2020 IEEE 39th International Performance Computing and Communications Conference (IPCCC)
影响因子:
--
作者:
[Danlin Jia;M. Saha;J. Bhimani;N. Mi]
通讯作者:
Danlin Jia;M. Saha;J. Bhimani;N. Mi
Collaborative Research: CNS core: OAC core: Small: New Techniques for I/O Behavior Modeling and Persistent Storage Device Configuration
-
批准号:2008072
-
项目类别:Standard Grant
-
资助金额:$24.49万
-
财政年份:2020
-
负责人:Ningfang Mi
-
依托单位:
CSR: EAGER: An Integrated Framework for Performance and Reliability in Large-scaled Computing Systems
-
批准号:1251129
-
项目类别:Standard Grant
-
资助金额:$27.24万
-
财政年份:2012
-
负责人:Ningfang Mi
-
依托单位:
海外基金