Using Model Trees for Computer Architecture Performance Analysis of Software Applications

Using Model Trees for Computer Architecture Performance Analysis of Software Applications
复制标题

DOI:
10.1109/ispass.2007.363742
复制
发表时间:
2007-04
期刊:
2007 IEEE International Symposium on Performance Analysis of Systems & Software
影响因子:
--
通讯作者:
ElMoustapha Ould-Ahmed-Vall;J. Woodlee;Charles R. Yount;K. Doshi;S. Abraham
ElMoustapha Ould-Ahmed-Vall;J. Woodlee;Charles R. Yount;K. Doshi;S. Abraham
中科院分区:
其他
文献类型:
--
作者:
ElMoustapha Ould-Ahmed-Vall;J. Woodlee;Charles R. Yount;K. Doshi;S. Abraham

文献摘要

被引文献

相似文献

识别特定计算机体系结构上的性能问题具有各种重要的好处,例如调整软件以提高性能,比较各种平台的性能以及协助设计新平台。为了实现这种分析,大多数现代微处理器提供对基于硬件的事件计数器的访问。不幸的是,无序执行、预取和推测等特性使原始数据的解释变得复杂。因此,为每个事件分配统一的估计惩罚的传统方法不能准确地识别和量化性能限制因素。本文提出了一种新的方法,采用统计回归建模方法,以更好地实现这一目标。具体来说,一个基于M5'算法的模型树的方法的基础上实现和验证,占事件的相互作用和工作负载的特点。该算法使用来自SPEC CPU 2006套件子集的数据自动构建性能模型树,识别套件中的唯一性能类(阶段),并将每个类与性能事件的唯一解释性线性模型相关联。这些模型可用于识别给定工作负载的性能问题,并估计解决每个问题的潜在收益。这些信息可以帮助确定性能优化工作的方向,将可用的时间和资源集中在最有可能影响性能问题并具有最高潜在收益的技术上。该模型树具有较高的相关性(大于0.98)和较低的相对绝对误差(小于8%),证明它是一种现代超标量电机性能分析的有效方法
The identification of performance issues on specific computer architectures has a variety of important benefits such as tuning software to improve performance, comparing the performance of various platforms and assisting in the design of new platforms. In order to enable this analysis, most modern micro-processors provide access to hardware-based event counters. Unfortunately, features such as out-of-order execution, pre-fetching and speculation complicate the interpretation of the raw data. Thus, the traditional approach of assigning a uniform estimated penalty to each event does not accurately identify and quantify performance limiters. This paper presents a novel method employing a statistical regression-modeling approach to better achieve this goal. Specifically, a model-tree based approach based on the M5' algorithm is implemented and validated that accounts for event interactions and workload characteristics. Data from a subset of the SPEC CPU2006 suite is used by the algorithm to automatically build a performance-model tree, identifying the unique performance classes (phases) found in the suite and associating with each class a unique, explanatory linear model of performance events. These models can be used to identify performance problems for a given workload and estimate the potential gain from addressing each problem. This information can help orient the performance optimization efforts to focus available time and resources on techniques most likely to impact performance problems with highest potential gain. The model tree exhibits high correlation (more than 0.98) and low relative absolute error (less than 8 %) between predicted and measured performance, attesting it as a sound approach for performance analysis of modern superscalar machines