PIMCloud: QoS-Aware Resource Management of Latency-Critical Applications in Clouds with Processing-in-Memory

PIMCloud: QoS-Aware Resource Management of Latency-Critical Applications in Clouds with Processing-in-Memory
复制标题

DOI:
10.1109/hpca53966.2022.00083
复制
发表时间:
2022-04
期刊:
2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Shuang Chen;Yi Jiang;Christina Delimitrou;José F. Martínez
Shuang Chen;Yi Jiang;Christina Delimitrou;José F. Martínez
中科院分区:
其他
文献类型:
--
作者:
Shuang Chen;Yi Jiang;Christina Delimitrou;José F. Martínez

文献摘要

相似文献

摩尔定律的放缓,加上逻辑和记忆的3D堆叠的进步,促使建筑师重新审视了内存处理概念(PIM),以克服记忆墙瓶颈。二十年前的景观随着越来越多的计算转移到云的转移。在关键延迟应用程序的上下文中,我们采用PIM架构。 PIM向这些服务展示的机会忽略了,并显示了正确管理与PIM相关的资源以满足交互式服务的QoS目标的重要性,然后我们提出了PIMCLOUD,这是一种QoS-Aware的资源经理,这是一家QoS-Aware Resource Manager,该资源经理使用QOS-ISAL SAMPARY设计了云系统。 PIM允许多个延迟至关重要和最佳效率应用托管。 APP和3-APP混合,与最先进的经理相比;(2)可帮助延迟式应用程序符合QoS;(3)适应不同的负载模式。
The slowdown of Moore’s Law, combined with advances in 3D stacking of logic and memory, have pushed architects to revisit the concept of processing-in-memory (PIM) to overcome the memory wall bottleneck. This PIM renaissance finds itself in a very different computing landscape from the one twenty years ago, as more and more computation shifts to the cloud. Most PIM architecture papers still focus on best-effort applications, while PIM’s impact on latency-critical cloud applications is not well understood.This paper explores how datacenters can exploit PIM architectures in the context of latency-critical applications. We adopt a general-purpose cloud server with HBM-based, 3D-stacked logic+memory modules, and study the impact of PIM on six diverse interactive cloud applications. We reveal the previously neglected opportunity that PIM presents to these services, and show the importance of properly managing PIM-related resources to meet the QoS targets of interactive services and maximize resource efficiency. Then, we present PIMCloud, a QoS-aware resource manager designed for cloud systems with PIM allowing colocation of multiple latency-critical and best-effort applications. We show that PIMCloud efficiently manages PIM resources: it (1) improves effective machine utilization by up to 70% and 85% (average 24% and 33%) under 2-app and 3-app mixes, compared to the best state-of-the-art manager; (2) helps latency-critical applications meet QoS; and (3) adapts to varying load patterns.