Quantitative Evaluation of Intel PEBS Overhead for Online System-Noise Analysis

Quantitative Evaluation of Intel PEBS Overhead for Online System-Noise Analysis
复制标题

DOI:
10.1145/3095770.3095773
复制
发表时间:
2017-06
期刊:
Proceedings of the 7th International Workshop on Runtime and Operating Systems for Supercomputers ROSS 2017
影响因子:
--
通讯作者:
Soramichi Akiyama;Takahiro Hirofuchi
Soramichi Akiyama;Takahiro Hirofuchi
中科院分区:
其他
文献类型:
--
作者:
Soramichi Akiyama;Takahiro Hirofuchi

文献摘要

相似文献

分析高通量系统引起的系统噪声(例如,Spark、RDBMS)必须在消息或请求级别的粒度上找到性能异常的根本原因,因为消息在很短的时间内通过许多组件传递。为此,我们认为使用配备在英特尔CPU中的基于事件的精确采样(PEBS),以比通常使用的更高的采样率是有希望的。它保存上下文信息(例如,通用寄存器)在发生各种硬件事件(例如高速缓存未命中)时。该信息可用于将由系统噪声引起的性能异常与特定消息相关联。一个挑战是尚未研究具有高采样率的PEBS开销的定量分析。这一点很关键,因为高采样率可能会导致严重的开销,但性能问题通常只在真实的环境中重现。在本文中,我们评估了PEBS的开销,并显示:(1)每次PEBS保存上下文信息时,由于PEBS的CPU开销,目标工作负载减慢200-300 ns,(2)CPU开销可以用于高精度地预测复杂工作负载(包括多线程工作负载)引起的实际开销,(3)由于PEBS将数据写入CPU缓存,因此PEBS会导致缓存污染和额外的内存IO,缓存污染的严重程度受采样率和分配给PEBS的缓存大小的影响。据我们所知,我们是第一个定量分析PEBS的开销。
Analyzing system-noise incurred to high-throughput systems (e.g., Spark, RDBMS) from the underlying machines must be in the granularity of the message- or request-level to find the root causes of performance anomalies, because messages are passed through many components in very short periods. To this end, we consider using Precise Event Based Sampling (PEBS) equipped in Intel CPUs at higher sampling rates than used normally is promising. It saves context information (e.g., the general purpose registers) at occurrences of various hardware events such as cache misses. The information can be used to associate performance anomalies caused by system noise with specific messages. One challenge is that quantitative analysis of PEBS overhead with high sampling rates has not yet been studied. This is critical because high sampling rates can cause severe overhead but performance problems are often reproducible only in real environments. In this paper, we evaluate the overhead of PEBS and show: (1) every time PEBS saves context information, the target workload slows down by 200-300 ns due to the CPU overhead of PEBS, (2) the CPU overhead can be used to predict actual overhead incurred with complex workloads including multi-threaded ones with high accuracy, and (3) PEBS incurs cache pollution and extra memory IO since PEBS writes data into the CPU cache, and the severity of cache pollution is affected both by the sampling rate and the buffer size allocated for PEBS. To the best of our knowledge, we are the first to quantitatively analyze the overhead of PEBS.