Improving data cache performance by pre-executing instructions under a cache miss

Improving data cache performance by pre-executing instructions under a cache miss
复制标题

通过在缓存未命中时预先执行指令来提高数据缓存性能

DOI:
--
复制
发表时间:
1997
期刊:
International Conference on Supercomputing
影响因子:
--
通讯作者:
T. Mudge
T. Mudge
中科院分区:
--
文献类型:
--
作者:
James Dundas;T. Mudge

文献摘要

被引文献

相似文献

PHiPAC是一种早期尝试,通过在可能的实现的大型设计空间中搜索以找到最佳方案来提高软件性能。在20世纪90年代初期,当时最有效的数值线性代数库是针对特定微架构和编译器精心手工调整的,并且通常用汇编语言编写。这使得算法能够针对当前平台的具体情况进行非常精确的调整,并为实现高效率提供了巨大机会。当时盛行的观点是,这种方法对于产生接近峰值的性能是必要的。另一方面,这种方法很脆弱,需要大量人力来尝试每种代码变体,因此只能探索可能的代码设计点的极小一部分。更糟糕的是,考虑到编译器和微架构的综合复杂性,很难预测哪些代码变体值得投入实现的努力。PHiPAC通过使用代码生成器规避了这种努力,这些代码生成器可以轻松地在一个设计空间内,甚至在完全不同的设计空间中生成各种各样非常不同的点。通过遵循一套精心制定的编码准则,生成的代码对于设计空间中的任何点都具有合理的效率。为了搜索设计空间,PHiPAC采取了一种相当简单但有效的方法。由于计算系统是由人设计且具有确定性的性质,人们可能会合理地认为,对微处理器和编译器进行智能建模足以在不进行任何计时的情况下预测给定算法的最佳点。但是优化编译器和动态调度微处理器的组合……
PHiPAC was an early attempt to improve software performance by searching in a large design space of possible implementations to find the best one. At the time, in the early 1990s, the most efficient numerical linear algebra libraries were carefully hand tuned for specific microarchitectures and compilers, and were often written in assembly language. This allowed very precise tuning of an algorithm to the specifics of a current platform, and provided great opportunity for high efficiency. The prevailing thought at the time was that such an approach was necessary to produce near-peak performance. On the other hand, this approach was brittle, and required great human effort to try each code variant, and so only a tiny subset of the possible code design points could be explored. Worse, given the combined complexities of the compiler and microarchitecture, it was difficult to predict which code variants would be worth the implementation effort. PHiPAC circumvented this effort by using code generators that could easily generate a vast assortment of very different points within a design space, and even across very different design spaces altogether. By following a set of carefully crafted coding guidelines, the generated code was reasonably efficient for any point in the design space. To search the design space, PHiPAC took a rather naive but effective approach. Due to the human-designed and deterministic nature of computing systems, one might reasonably think that smart modeling of the microprocessor and compiler would be sufficient to predict, without performing any timing, the optimal point for a given algorithm. But the combination of an optimizing compiler and a dynamically scheduled mi-