Datamime: Generating Representative Benchmarks by Automatically Synthesizing Datasets

Datamime: Generating Representative Benchmarks by Automatically Synthesizing Datasets
复制标题

DOI:
10.1109/micro56248.2022.00082
复制
发表时间:
2022-10
期刊:
2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Hyun Ryong Lee;Daniel Sánchez
Hyun Ryong Lee;Daniel Sánchez
中科院分区:
其他
文献类型:
--
作者:
Hyun Ryong Lee;Daniel Sánchez

文献摘要

相似文献

与生产工作负载的行为非常匹配的基准测试对于设计和提供计算机系统至关重要。但是,当前的方法不足:首先,开源基准使用公共数据集,这些数据集导致生产工作负载不同的行为。其次,BlackBox Workload克隆技术生成了模仿目标工作负载的合成代码,但是结果程序未能捕获大多数工作负载特征,例如微体系结构瓶颈或随时间变化的行为。模拟复杂应用程序的生成代码是一个极其困难的问题。相反,我们提出了一种基准合成的不同,更轻松的方法。我们的主要见解是,对于许多生产工作负载,该程序已公开可用,或者有一个相似的开源程序。在这种情况下,生成正确的数据集足以产生准确的基准测试。基于此观察结果,我们提出了Datamime,这是一种为生产工作负载生成代表性基准的配置文件引导的方法。 Datamime使用目标工作负载的性能配置文件来生成数据集,该数据集在基准程序中使用时的行为与目标工作负载非常相似。 Datamime生成的合成基准测试与这些工作负载的微体系特征非常匹配,而IPC的平均绝对百分比误差为3.2%。微构造行为在处理器类型之间保持近距离。最后,还复制了时间变化的行为,使这些基准对例如表征并优化尾部潜伏期。
Benchmarks that closely match the behavior of production workloads are crucial to design and provision computer systems. However, current approaches fall short: First, open-source benchmarks use public datasets that cause different behavior from production workloads. Second, blackbox workload cloning techniques generate synthetic code that imitates the target workload, but the resulting program fails to capture most workload characteristics, such as microarchitectural bottlenecks or time-varying behavior.Generating code that mimics a complex application is an extremely hard problem. Instead, we propose a different and easier approach to benchmark synthesis. Our key insight is that, for many production workloads, the program is publicly available or there is a reasonably similar open-source program. In this case, generating the right dataset is sufficient to produce an accurate benchmark.Based on this observation, we present Datamime, a profile-guided approach to generate representative benchmarks for production workloads. Datamime uses the performance profiles of a target workload to generate a dataset that, when used by a benchmark program, behaves very similarly to the target workload in terms of its microarchitectural characteristics.We evaluate Datamime on several datacenter workloads. Datamime generates synthetic benchmarks that closely match the microarchitectural features of these workloads, with a mean absolute percentage error of 3.2% on IPC. Microarchitectural behavior stays close across processor types. Finally, time-varying behaviors are also replicated, making these benchmarks useful to e.g. characterize and optimize tail latency.