An automated framework for characterizing and subsetting GPGPU workloads

An automated framework for characterizing and subsetting GPGPU workloads
复制标题

用于表征和子集 GPGPU 工作负载的自动化框架

DOI:
10.1109/ispass.2016.7482105
复制
发表时间:
2016
期刊:
2016 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)
影响因子:
--
通讯作者:
Wu
Wu
中科院分区:
--
文献类型:
--
作者:
Vignesh Adhinarayanan;Wu

文献摘要

被引文献

相似文献

图形处理单元(GPU)由于其相对于其成本的上级性能和能量效率而在当今的计算系统中变得越来越普遍。为了进一步改善这些所需的特性,研究人员提出了几种软件和硬件技术。这些建议的技术的评估可能是棘手的,由于特设的性质,其中的应用程序被选择进行评估。有时研究人员会花费不必要的时间来评估冗余的工作负载,这对于涉及仿真的耗时研究来说尤其成问题。其他时候,当选择太少的工作负载进行评估时,他们无法暴露他们提出的技术的缺点。为了克服这些问题,我们提出了一个自动化的框架,根据用户选择的一组性能指标/计数器的特点和子集GPGPU的工作负载。该框架在内部使用主成分分析(PCA)来降低所选指标的维度,然后使用层次聚类来识别工作负载之间的相似性。在这项研究中,我们使用我们的框架,以确定冗余在最近发布的SPEC ACCEL OpenCL基准套件使用一些架构相关的指标。我们的分析表明,在19个应用程序的基准测试套件中,8个应用程序的子集提供了大部分的多样性。我们还对Parboil、Rodinia和SHOC基准套件进行了子集化,然后将它们相互比较,以确定这些套件中的“差距”。作为一个例子,我们表明,SHOC有许多应用程序是彼此相似的,可以受益于从Parboil增加四个应用程序,以提高其多样性。
Graphics processing units (GPUs) are becoming increasingly common in today's computing systems due to their superior performance and energy efficiency relative to their cost. To further improve these desired characteristics, researchers have proposed several software and hardware techniques. Evaluation of these proposed techniques could be tricky due to the ad-hoc nature in which applications are selected for evaluation. Sometimes researchers spend unnecessary time evaluating redundant workloads, which is particularly problematic for time-consuming studies involving simulation. Other times, they fail to expose the shortcomings of their proposed techniques when too few workloads are chosen for evaluation. To overcome these problems, we propose an automated framework that characterizes and subsets GPGPU workloads, depending on a user-chosen set of performance metrics/counters. This framework internally uses principal component analysis (PCA) to reduce the dimensionality of the chosen metrics and then uses hierarchical clustering to identify similarity among the workloads. In this study, we use our framework to identify redundancy in the recently released SPEC ACCEL OpenCL benchmark suite using a few architecture-dependent metrics. Our analysis shows that a subset of eight applications provides most of the diversity in the 19-application benchmark suite. We also subset the Parboil, Rodinia, and SHOC benchmark suites and then compare them against each other to identify “gaps” in these suites. As an example, we show that SHOC has many applications that are similar to each other and could benefit from adding four applications from Parboil to improve its diversity.