CAREER: Capacity Planning Methodologies for Large Clusters with Heterogeneous Architectures and Diverse Applications
CAREER: Capacity Planning Methodologies for Large Clusters with Heterogeneous Architectures and Diverse Applications
批准号:
1452751
负责人:
Ningfang Mi
金额:
$45.96万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-04-01 至 2022-03-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
This project focuses on developing innovative techniques and algorithms to build adequate system models and support performance and reliability analysis in order to better explain large system behavior, predict application performance, and ensure high resource efficiency and system dependability. Large cluster environments are an important part of today's computing infrastructure, providing the platform for running applications that handle core business and operational data. However, with the complexity of computing and application infrastructure increasing and the requirements for high quality of service growing, large cluster environments are facing the difficult task of ensuring that applications are always available and delivering adequate performance. This project expects to achieve new capacity planning techniques for performance modeling, workload measurements and model parameterizations of large cluster systems. Intelligent capacity and reliability modeling will enable service providers to determine the best platform for their application before deploying and running the application. It will also enable system managers to optimize the performance, reliability and efficiency of the entire cluster infrastructure. This research will develop new performance modeling methods to capture the characteristics of heterogeneous hardware architectures and predict the behavior of an application running on an array of computing platforms. The research will extend performance modeling to failure awareness. The improved models will enable an accurate prediction of performance and reliability of a complex large-scale system by capturing the characteristics of both system workloads and failure events. In addition, the researchers will develop new advanced techniques to parameterize performance models with essential processing information of computational and communication components. These essential processing information do not only limit to mean values but also include other critical yet complicated features such as resource contention and burstiness symptoms. The project is involved with educational activities reaching out to students from secondary to graduate schools to aggressively motivate students, especially women, towards science and engineering integrating this research into curriculum development and undergraduate research activities.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1109/ipccc50635.2020.9391566
发表时间:
2020-11
期刊:
2020 IEEE 39th International Performance Computing and Communications Conference (IPCCC)
影响因子:
--
作者:
[Danlin Jia;M. Saha;J. Bhimani;N. Mi]
通讯作者:
Danlin Jia;M. Saha;J. Bhimani;N. Mi
Collaborative Research: CNS core: OAC core: Small: New Techniques for I/O Behavior Modeling and Persistent Storage Device Configuration
-
批准号:2008072
-
项目类别:Standard Grant
-
资助金额:$24.49万
-
财政年份:2020
-
负责人:Ningfang Mi
-
依托单位:
CSR: EAGER: An Integrated Framework for Performance and Reliability in Large-scaled Computing Systems
-
批准号:1251129
-
项目类别:Standard Grant
-
资助金额:$27.24万
-
财政年份:2012
-
负责人:Ningfang Mi
-
依托单位:
海外基金