DR-BW: Identifying Bandwidth Contention in NUMA Architectures with Supervised Learning

DR-BW: Identifying Bandwidth Contention in NUMA Architectures with Supervised Learning
复制标题

DR-BW:通过监督学习识别 NUMA 架构中的带宽争用

DOI:
10.1109/ipdps.2017.97
复制
发表时间:
2017
期刊:
2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Xu Liu
Xu Liu
中科院分区:
--
文献类型:
--
作者:
Hao Xu;Shasha Wen;Alfredo Giménez;T. Gamblin;Xu Liu

文献摘要

被引文献

相似文献

非均匀存储器访问(NUMA)体系结构被广泛应用于主流多插槽计算机系统中以扩展存储器带宽。如果没有NUMA感知设计,程序可能会因插座间带宽争用而显著降低性能。然而,识别带宽争用是一项具有挑战性的工作。现有方法测量带宽消耗。然而,单靠消耗还不足以量化带宽争用。此外,现有方法诊断整个程序执行的带宽,但缺乏将带宽性能与所涉及的源代码和数据结构相关联的能力。为了应对这些挑战,我们提出了DR-BW,这是一个基于机器学习的新工具,用于识别NUMA体系结构中的带宽竞争并提供优化指导。DR-BW首先训练一组微观基准,并通过有监督的机器学习模型提取有用的特征来识别带宽竞争。实验表明,DR-BW算法达到了96%以上的准确率。其次,DR-BW将引起带宽竞争的内存访问与数据对象相关联,这为优化提供了直观的指导。第三,我们将DR-BW应用于许多实际基准。我们的优化基于从DR-BW获得的见解,在现代NUMA体系结构中的加速比高达6.5倍。
Non-Uniform Memory Access (NUMA) architectures are widely used in mainstream multi-socket computer systems to scale memory bandwidth. Without a NUMA-aware design, programs can suffer from significant performance degradation due to inter-socket bandwidth contention. However, identifying bandwidth contention is challenging. Existing methods measure bandwidth consumption. However, consumption alone is insufficient to quantify bandwidth contention. Furthermore, existing methods diagnose bandwidth for the entire program execution, but lack the ability to associate bandwidth performance to the source code and data structures involved. To address these challenges, we propose DR-BW, a new tool based on machine learning to identify bandwidth contention in NUMA architectures and provide optimization guidance. DR-BW first trains a set of micro benchmarks and extracts useful features to identify bandwidth contention via a supervised machine learning model. Our experiments show that DR-BW achieves more than 96% accuracy. Second, DR-BW associates memory accesses that incur bandwidth contention with data objects, which provides intuitive guidance for optimization. Third, we apply DR-BW to a number of real benchmarks. Our optimization based on the insights obtained from DR-BW yields up to a 6.5× speedup in modern NUMA architectures.