ILP and TLP in shared memory applications: A limit study

ILP and TLP in shared memory applications: A limit study
复制标题

共享内存应用中的 ILP 和 TLP:极限研究

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
Paul V. Gratz
Paul V. Gratz
中科院分区:
--
文献类型:
--
作者:
Ehsan Fatehi;Paul V. Gratz

文献摘要

被引文献

相似文献

随着Dennard缩放的崩溃,随着芯片多处理器(CMP)设计向多核扩展,未来的处理器设计将受到功率限制的影响。因此,未来的CMP在功率性能效率方面进行优化设计至关重要。必须对未来工作负载进行特性分析,以确保每瓦消耗的性能回报最大化。因此,有必要对新兴工作负载进行详细分析,以了解它们在功耗和性能权衡方面相对于硬件的特性。在本文中,我们进行了限制研究,同时分析了两种主要形式的并行利用现代计算机体系结构:指令级并行(ILP)和线程级并行(TLP)。这项研究提供了洞察未来架构可以实现的性能上限。此外,它还确定了新兴工作负载的瓶颈。据我们所知,我们的工作是第一个研究,结合两种形式的并行到一个研究与现代应用。我们评估PARSEC多线程基准套件使用专门的跟踪驱动的模拟器。我们做出了一些贡献,描述了下一代应用程序的高级行为。例如,我们发现这些应用程序包含的ILP比当前从真实的机器中提取的ILP多929倍。然后,我们打破了应用程序的线程数量不断增加(利用TLP),指令窗口大小,现实的分支预测,现实的内存延迟,线程依赖于可利用的ILP的影响。我们的研究表明,这些基准彼此差异很大。因此,我们预计没有单一的,同质的,微架构将最佳地为所有工作,主张可重构的,异构的设计。
With the breakdown of Dennard scaling, future processor designs will be at the mercy of power limits as Chip MultiProcessor (CMP) designs scale out to many-cores. It is critical, therefore, that future CMPs be optimally designed in terms of performance efficiency with respect to power. A characterization analysis of future workloads is imperative to ensure maximum returns of performance per Watt consumed. Hence, a detailed analysis of emerging workloads is necessary to understand their characteristics with respect to hardware in terms of power and performance tradeoffs. In this paper, we conduct a limit study simultaneously analyzing the two dominant forms of parallelism exploited by modern computer architectures: Instruction Level Parallelism (ILP) and Thread Level Parallelism (TLP). This study gives insights into the upper bounds of performance that future architectures can achieve. Furthermore it identifies the bottlenecks of emerging workloads. To the best of our knowledge, our work is the first study that combines the two forms of parallelism into one study with modern applications. We evaluate the PARSEC multithreaded benchmark suite using a specialized trace-driven simulator. We make several contributions describing the high-level behavior of next-generation applications. For example, we show these applications contain up to a factor of 929× more ILP than what is currently being extracted from real machines. We then show the effects of breaking the application into increasing numbers of threads (exploiting TLP), instruction window size, realistic branch prediction, realistic memory latency, and thread dependencies on exploitable ILP. Our examination shows that theses benchmarks differed vastly from one another. As a result, we expect no single, homogeneous, micro-architecture will work optimally for all, arguing for reconfigurable, heterogeneous designs.