Pythia: Improving Datacenter Utilization via Precise Contention Prediction for Multiple Co-located Workloads

Pythia: Improving Datacenter Utilization via Precise Contention Prediction for Multiple Co-located Workloads
复制标题

DOI:
10.1145/3274808.3274820
复制
发表时间:
2018-11
期刊:
Proceedings of the 19th International Middleware Conference
影响因子:
--
通讯作者:
Ran Xu;S. Mitra;Jason Rahman;Peter Bai;Bowen Zhou;G. Bronevetsky;S. Bagchi
Ran Xu;S. Mitra;Jason Rahman;Peter Bai;Bowen Zhou;G. Bronevetsky;S. Bagchi
中科院分区:
其他
文献类型:
--
作者:
Ran Xu;S. Mitra;Jason Rahman;Peter Bai;Bowen Zhou;G. Bronevetsky;S. Bagchi

文献摘要

被引文献

相似文献

随着现代体系结构中核心数量的增加,共同关注多个工作负载的需求对于改善整体计算利用率至关重要。但是,通常避免在同一服务器上共同关注多个工作负载,以保护延迟敏感(LS)工作负载的性能免受其他共享资源上其他共同确定的工作负载(例如缓存和内存带宽)上的争论。在本文中,我们介绍了毕田(Pythia),这是一位共同定位的经理,当多个共同确定的工作负载干扰LS工作负载时,可以精确预测共享资源的合并争议。毕田(Pythia)使用一个简单的线性回归模型,该模型可以使用所有可能的共线的大型配置空间的一小部分训练,并且仍然可以对合并论点做出高度准确的预测。基于这些预测,毕达斯明智地安排了输入的工作量,以便在不违反LS工作负载的延迟阈值的情况下改善集群利用率。我们证明,与当用户准备在QoS指标中牺牲高达5%并实现99%的群集利用时,毕田的调度可以提高群集利用率71%,而如果QoS中的10%降级是QoS中的10%降级,则可以将群集利用率提高71%。可以接受。
With the increase in the number cores in modern architectures, the need for co-locating multiple workloads has become crucial for improving the overall compute utilization. However, co-locating multiple workloads on the same server is often avoided to protect the performance of the latency sensitive (LS) workloads from the contentions created by other co-located workloads on the shared resources, such as cache and memory bandwidth. In this paper, we present Pythia, a co-location manager that can precisely predict the combined contention on shared resources when multiple co-located workloads interfere with an LS workload. Pythia uses a simple linear regression model that can be trained using a small fraction of the large configuration space of all possible co-locations and can still make highly accurate predictions for the combined contentions. Based on those predictions, Pythia judiciously schedules incoming workloads so that cluster utilization is improved without violating the latency threshold of the LS workloads. We demonstrate that Pythia's scheduling can improve cluster utilization by 71% compared to a simple extension of a prior work when the user is ready to sacrifice up to 5% in the QoS metric and achieve cluster utilization of 99% if 10% degradation in QoS is acceptable.