Rigorous benchmarking in reasonable time

Rigorous benchmarking in reasonable time
复制标题

DOI:
10.1145/2464157.2464160
复制
发表时间:
2013-06
期刊:
--
影响因子:
--
通讯作者:
T. Kalibera;Richard E. Jones
T. Kalibera;Richard E. Jones
中科院分区:
其他
文献类型:
--
作者:
T. Kalibera;Richard E. Jones

文献摘要

被引文献

相似文献

实验评估是系统研究的关键。由于现代系统是复杂且非确定性的,因此良好的实验方法要求研究人员解释不确定性。为了获得有效的结果,预计他们将多次使用基准的许多迭代,调用虚拟机(VM),甚至不止一次重建VM或基准二进制文件。所有这些重复花费时间完成实验。当前,许多评估放弃了足够的重复或严格的统计方法,甚至仅在训练尺寸中运行基准。报道的结果通常缺乏适当的变化估计值,当报告两个系统之间的微小差异时,有些简直不可靠。相反,我们为重复和总结结果提供了一种有效利用实验时间的结果的统计严格方法。时间效率来自两个关键观察。首先,在给定平台上的给定基准通常容易出现的非确定性要比常见的发表角案例研究的常见案例少得多。其次,在发生大多数不确定性的地方最需要重复(无论是在构建,执行之间还是在迭代之间)。我们使用一种新型的数学模型来捕获实验成本,我们用来确定必要且足够的实验级别的重复数量,以获得给定的精度。我们将方法介绍为一本食谱,可指导研究人员应进行的重复数量以获得可靠的结果。我们还展示了如何使用效果大小置信区间提出结果。例如,我们展示了如何使用我们的方法在三个最近的平台上使用DACAPO和SPEC CPU基准进行吞吐量实验。
Experimental evaluation is key to systems research. Because modern systems are complex and non-deterministic, good experimental methodology demands that researchers account for uncertainty. To obtain valid results, they are expected to run many iterations of benchmarks, invoke virtual machines (VMs) several times, or even rebuild VM or benchmark binaries more than once. All this repetition costs time to complete experiments. Currently, many evaluations give up on sufficient repetition or rigorous statistical methods, or even run benchmarks only in training sizes. The results reported often lack proper variation estimates and, when a small difference between two systems is reported, some are simply unreliable. In contrast, we provide a statistically rigorous methodology for repetition and summarising results that makes efficient use of experimentation time. Time efficiency comes from two key observations. First, a given benchmark on a given platform is typically prone to much less non-determinism than the common worst-case of published corner-case studies. Second, repetition is most needed where most uncertainty arises (whether between builds, between executions or between iterations). We capture experimentation cost with a novel mathematical model, which we use to identify the number of repetitions at each level of an experiment necessary and sufficient to obtain a given level of precision. We present our methodology as a cookbook that guides researchers on the number of repetitions they should run to obtain reliable results. We also show how to present results with an effect size confidence interval. As an example, we show how to use our methodology to conduct throughput experiments with the DaCapo and SPEC CPU benchmarks on three recent platforms.