Synthesizing benchmarks for predictive modeling

Synthesizing benchmarks for predictive modeling
复制标题

DOI:
10.1109/cgo.2017.7863731
复制
发表时间:
2017-02
影响因子:
10.9
通讯作者:
Chris Cummins;Pavlos Petoumenos;Zheng Wang-;Hugh Leather
Chris Cummins;Pavlos Petoumenos;Zheng Wang-;Hugh Leather
中科院分区:
工程技术1区
文献类型:
--
作者:
Chris Cummins;Pavlos Petoumenos;Zheng Wang-;Hugh Leather

文献摘要

相似文献

基于机器学习的预测建模是构建编译器性能的有效方法,但缺乏测试基准。编译领域之外的典型机器学习实验会训练数千或数百万个示例。然而,在编译器的机器学习中,通常只有几十个常见的基准测试可用。这限制了学习模型的质量,因为它们通常具有非常稀疏的高维特征空间的训练数据。我们需要的是一种生成无限数量的训练程序的方法,这些训练程序可以精细地覆盖特征空间。同时,生成的程序必须与人类开发人员实际编写的程序类型相似,否则学习将针对特征空间的错误部分。我们挖掘开源存储库中的程序片段,并应用深度学习技术自动构建人类如何编写程序的模型。我们对这些模型进行采样,以生成无限数量的可运行训练程序。程序的质量是这样的,即使是人类开发人员也很难区分我们生成的程序和手写的代码。我们使用我们的OpenCL程序生成器CLgen来自动合成数千个程序,并表明对这些程序的学习将最先进的预测模型的性能提高了1.27%。此外,特征空间的精细覆盖自动暴露了特征设计中的弱点,这些弱点在现有基准测试套件的稀疏训练示例中是不可见的。纠正这些弱点进一步提高了4.30英寸的性能。
Predictive modeling using machine learning is an effective method for building compiler heuristics, but there is a shortage of benchmarks. Typical machine learning experiments outside of the compilation field train over thousands or millions of examples. In machine learning for compilers, however, there are typically only a few dozen common benchmarks available. This limits the quality of learned models, as they have very sparse training data for what are often high-dimensional feature spaces. What is needed is a way to generate an unbounded number of training programs that finely cover the feature space. At the same time the generated programs must be similar to the types of programs that human developers actually write, otherwise the learning will target the wrong parts of the feature space. We mine open source repositories for program fragments and apply deep learning techniques to automatically construct models for how humans write programs. We sample these models to generate an unbounded number of runnable training programs. The quality of the programs is such that even human developers struggle to distinguish our generated programs from hand-written code. We use our generator for OpenCL programs, CLgen, to automatically synthesize thousands of programs and show that learning over these improves the performance of a state of the art predictive model by 1.27� . In addition, the fine covering of the feature space automatically exposes weaknesses in the feature design which are invisible with the sparse training examples from existing benchmark suites. Correcting these weaknesses further increases performance by 4.30� .