The impact of memory subsystem resource sharing on datacenter applications

The impact of memory subsystem resource sharing on datacenter applications
复制标题

DOI:
10.1145/2000064.2000099
复制
发表时间:
2011-06
期刊:
2011 38th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Lingjia Tang;Jason Mars;Neil Vachharajani;R. Hundt;M. Soffa
Lingjia Tang;Jason Mars;Neil Vachharajani;R. Hundt;M. Soffa
中科院分区:
其他
文献类型:
--
作者:
Lingjia Tang;Jason Mars;Neil Vachharajani;R. Hundt;M. Soffa

文献摘要

被引文献

相似文献

在本文中,我们研究了共享内存资源对五个 Google 数据中心应用程序的影响:网络搜索引擎、bigtable、内容分析器、图像拼接和协议缓冲区。虽然之前的工作没有发现跨 PARSEC 基准套件的缓存共享产生积极或消极影响,但我们发现,在这些数据中心应用程序中,不正确地共享资源既带来了相当大的好处,也带来了潜在的退化。在本文中,我们首先研究了数据中心应用程序的线程到核心映射的重要性,因为线程可以映射为共享或不共享缓存和总线带宽。其次,我们研究了具有不同内存行为的多个应用程序的共置线程的影响,并发现给定应用程序的最佳映射会根据其共同运行者而变化。第三,我们研究了在各种线程到核心映射场景中影响性能的应用程序特征。最后,我们提出了一种基于启发式和自适应的方法,以在数据中心中做出良好的线程到核心决策。仅根据应用程序线程映射到内核的方式,我们观察到网络搜索的性能波动高达 25%,其他关键应用程序的性能波动高达 40%。通过采用我们的自适应线程到核心映射器,本工作中介绍的数据中心应用程序的性能比现状线程到核心映射提高了 22%,并且性能与最佳性能相差不到 3%。
In this paper we study the impact of sharing memory resources on five Google datacenter applications: a web search engine, bigtable, content analyzer, image stitching, and protocol buffer. While prior work has found neither positive nor negative effects from cache sharing across the PARSEC benchmark suite, we find that across these datacenter applications, there is both a sizable benefit and a potential degradation from improperly sharing resources. In this paper, we first present a study of the importance of thread-to-core mappings for applications in the datacenter as threads can be mapped to share or to not share caches and bus bandwidth. Second, we investigate the impact of co-locating threads from multiple applications with diverse memory behavior and discover that the best mapping for a given application changes depending on its co-runner. Third, we investigate the application characteristics that impact performance in the various thread-to-core mapping scenarios. Finally, we present both a heuristics-based and an adaptive approach to arrive at good thread-to-core decisions in the datacenter. We observe performance swings of up to 25% for web search and 40% for other key applications, simply based on how application threads are mapped to cores. By employing our adaptive thread-to-core mapper, the performance of the datacenter applications presented in this work improved by up to 22% over status quo thread-to-core mapping and performs within 3% of optimal.