Improving the Utilization of Micro-operation Caches in x86 Processors

Improving the Utilization of Micro-operation Caches in x86 Processors
复制标题

提高x86处理器中微操作缓存的利用率

DOI:
10.1109/micro50266.2020.00025
复制
发表时间:
2020
期刊:
2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
J. Kalamatianos
J. Kalamatianos
中科院分区:
--
文献类型:
--
作者:
Jagadish B. Kotra;J. Kalamatianos

文献摘要

参考文献

被引文献

相似文献

大多数现代处理器的员工可变长度,复杂的指令集计算(CISC)指令以获取能量成本和带宽的要求。 - 运行(也称为UOPS)。由UOP缓存派遣,(2)较低的解码器能量消耗,(3)早期检测到错误预测的分支。在本文中,我们观察到在某些UOP CACHE进入构造规则下,UOP缓存可以大大碎片。 ,我们提出了两个完整的优化来解决碎片:高速缓存边界无关UOP缓存设计(clasp)和UOP缓存压实。将单个UOP缓存进入。提高性能高达5.6%,当CLASP与最激进的压实变体相结合时,解码器的功率最高为19.63%,性能提高了12.8%,解码器的功率节省最高31.53%。
Most modern processors employ variable length, Complex Instruction Set Computing (CISC) instructions to reduce instruction fetch energy cost and bandwidth requirements. High throughput decoding of CISC instructions requires energy hungry logic for instruction identification. Efficient CISC instruction execution motivated mapping them to fixed length micro-operations (also known as uops). To reduce costly decoder activity, commercial CISC processors employ a micro-operations cache (uop cache) that caches uop sequences, bypassing the decoder. Uop cache’s benefits are: (1) shorter pipeline length for uops dispatched by the uop cache, (2) lower decoder energy consumption, and, (3) earlier detection of mispredicted branches.In this paper, we observe that a uop cache can be heavily fragmented under certain uop cache entry construction rules. Based on this observation, we propose two complementary optimizations to address fragmentation: Cache Line boundary AgnoStic uoP cache design (CLASP) and uop cache compaction. CLASP addresses the internal fragmentation caused by short, sequential uop sequences, terminated at the I-cache line boundary, by fusing them into a single uop cache entry. Compaction further lowers fragmentation by placing to the same uop cache entry temporally correlated, non-sequential uop sequences mapped to the same uop cache set. Our experiments on a x86 simulator using a wide variety of benchmarks show that CLASP improves performance up to 5.6% and lowers decoder power up to 19.63%. When CLASP is coupled with the most aggressive compaction variant, performance improves by up to 12.8% and decoder power savings are up to 31.53%.
调动微操作:利用上下文敏感解码实现安全性和能源效率
DOI: 10.1109/isca.2018.00058
发表时间: 2018
期刊: Intl Symposium on Computer Architecture
影响因子: --
作者:
Taram, Mohammadkazem;Venkat, Ashish;Tullsen, Dean
通讯作者: Tullsen, Dean