EAGER: Using Machine Learning to Increase the Operational Efficiency of Large Distributed Systems
EAGER: Using Machine Learning to Increase the Operational Efficiency of Large Distributed Systems
批准号:
1649087
负责人:
Evgenia Smirni
金额:
$30.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2019-08-31
中文摘要
大型分布式系统如今无处不在,是面向广泛客户和应用程序的可持续IT解决方案的一部分。私有云或公共云中的数据中心和高性能计算系统是复杂、高度分布式系统的两个例子:前者几乎每个人每天都在使用,后者被计算科学家用于推进科学和工程。这些复杂系统的高可用性和可靠性对用户体验的质量至关重要。这类系统的有效管理有助于它们的可用性和可靠性,并依赖于对用户集体需求的时间安排的先验知识和对各种系统组件的某些性能测量(例如,使用、温度、功率)的先验知识。该项目旨在提供一种系统的方法,通过开发能够在精细和粗略的时间尺度内高效和准确地预测到来的工作量的神经网络,来提高复杂的分布式系统的运行效率。这种工作负载预测可以通过推动专门旨在增强可靠性的主动式管理策略,显著提高数据中心和高性能系统的运营效率。对于数据中心,重点是积极减少由主动管理虚拟机调整大小和迁移而自动触发的性能票证。对于高性能计算系统,重点是预测硬件故障,以自动提高调度器的效率,直接冷却,并改善性能和内存带宽。
英文摘要
Large, distributed systems are nowadays ubiquitous and part of sustainable IT solutions to a broad range of customers and applications. Data centers in the private or public cloud and high performance computing systems are two examples of complex, highly distributed systems: the former are used by almost everyone on a daily basis, the latter are used by computational scientists for advancing science and engineering. High availability and reliability of these complex systems are important for the quality of user experience. Efficient management of such systems contributes to their availability and reliability, and relies on a priori knowledge of the timing of the collective demands of users and a priori knowledge of certain performance measures (e.g., usage, temperature, power) of various systems components.This project aims to provide a systematic methodology to improve the operational efficiency of complex, distributed systems by developing neural networks that can efficiently and accurately predict the incoming workload within fine and coarse time scales. Such workload prediction can dramatically improve the operational efficiency of data centers and high performance systems by driving proactive management strategies that specifically aim to enhance reliability. For datacenters, the focus is on actively reducing performance tickets that are automatically triggered by pro-actively managing virtual machine resizing and migration. For high performance computing systems the focus is on predicting hardware faults to autonomically improve the scheduler's efficiency, direct cooling, and improve performance and memory bandwidth.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/micro.2018.00066
发表时间:
2018-10
期刊:
2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
作者:
[Bin Nie;Lishan Yang;Adwait Jog;E. Smirni]
通讯作者:
Bin Nie;Lishan Yang;Adwait Jog;E. Smirni
DOI:
10.1109/tnsm.2018.2808352
发表时间:
2018-02
期刊:
IEEE Transactions on Network and Service Management
影响因子:
5.3
作者:
[Feng Yan;Yuxiong He;Olatunji Ruwase;E. Smirni]
通讯作者:
Feng Yan;Yuxiong He;Olatunji Ruwase;E. Smirni
DOI:
10.1109/cloud.2017.43
发表时间:
2017-06
期刊:
2017 IEEE 10th International Conference on Cloud Computing (CLOUD)
影响因子:
--
作者:
[Feng Yan;Lihua Ren;Daniel J. Dubois;G. Casale;Jiawei Wen;E. Smirni]
通讯作者:
Feng Yan;Lihua Ren;Daniel J. Dubois;G. Casale;Jiawei Wen;E. Smirni
CEDULE: A Scheduling Framework for Burstable Performance in Cloud Computing
CEDULE:云计算中突发性能的调度框架
DOI:
10.1109/icac.2018.00024
发表时间:
2018
期刊:
2018 IEEE International Conference on Autonomic Computing (ICAC
影响因子:
--
作者:
[Ali, Ahsan, Pinciroli, Riccardo, Yan, Feng, Smirni, Evgenia]
通讯作者:
Smirni, Evgenia
DOI:
10.1109/tnsm.2018.2794409
发表时间:
2018-01
期刊:
IEEE Transactions on Network and Service Management
影响因子:
5.3
作者:
[Ji Xue;R. Birke;L. Chen;E. Smirni]
通讯作者:
Ji Xue;R. Birke;L. Chen;E. Smirni
共 7 条
EAGER: Epidemic Spread Modeling Using Hard Data
-
批准号:2130681
-
项目类别:Standard Grant
-
资助金额:$20.57万
-
财政年份:2021
-
负责人:Evgenia Smirni
-
依托单位:
BIGDATA: IA: Collaborative Research: Protecting Yourself from Wildfire Smoke: Big Data-Driven Adaptive Air Quality Prediction Methodologies
-
批准号:1838022
-
项目类别:Standard Grant
-
资助金额:$29.83万
-
财政年份:2019
-
负责人:Evgenia Smirni
-
依托单位:
SHF-Small: Robust Methodologies for Effective Data Center Management
-
批准号:1218758
-
项目类别:Standard Grant
-
资助金额:$49.08万
-
财政年份:2012
-
负责人:Evgenia Smirni
-
依托单位:
CPA-ACR-CSA: Effective Resource Allocation under Temporal Dependence
-
批准号:0811417
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2008
-
负责人:Evgenia Smirni
-
依托单位:
CSR-SMA: Autocorrelated Flows in Systems: Analytic Models and Applications
-
批准号:0720699
-
项目类别:Continuing Grant
-
资助金额:$20.0万
-
财政年份:2007
-
负责人:Evgenia Smirni
-
依托单位:
ITR-(ASE)-(dmc+int): Reconfigurable, Data-driven Resource Allocation in Complex Systems: Practice and Theoretical Foundations
-
批准号:0428330
-
项目类别:Standard Grant
-
资助金额:$41.39万
-
财政年份:2004
-
负责人:Evgenia Smirni
-
依托单位:
Effective Techniques and Tools for Resource Management in Clustered Web Servers
-
批准号:0098278
-
项目类别:Continuing Grant
-
资助金额:$28.0万
-
财政年份:2001
-
负责人:Evgenia Smirni
-
依托单位:
Collaborative Research: Adaptive Data Parallel Storage
-
批准号:0090221
-
项目类别:Continuing Grant
-
资助金额:$18.54万
-
财政年份:2001
-
负责人:Evgenia Smirni
-
依托单位:
Next Generation Software: Coordinated Allocation of Processor and I/O Resources in Parallel Systems
-
批准号:9974992
-
项目类别:Continuing Grant
-
资助金额:$35.0万
-
财政年份:1999
-
负责人:Evgenia Smirni
-
依托单位:
国内基金
海外基金
Capture and Release of Droplets Using Advanced Materials for High Technology Applications
-
批准号:52073127
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2020
-
负责人:Alidad Amirfazli
-
依托单位:
Molecular Interaction Reconstruction of Rheumatoid Arthritis Therapies Using Clinical Data
-
批准号:31070748
-
项目类别:面上项目
-
资助金额:34.0万元
-
批准年份:2010
-
负责人:Christine Nardini
-
依托单位: