Optimizing Coherence Traffic in Manycore Processors Using Closed-Form Caching/Home Agent Mappings

Optimizing Coherence Traffic in Manycore Processors Using Closed-Form Caching/Home Agent Mappings
复制标题

DOI:
10.1109/access.2021.3058280
复制
发表时间:
2021
期刊:
影响因子:
3.9
通讯作者:
Steve Kommrusch;Marcos Horro;L. Pouchet;Gabriel Rodríguez;J. Touriño
Steve Kommrusch;Marcos Horro;L. Pouchet;Gabriel Rodríguez;J. Touriño
中科院分区:
计算机科学3区
文献类型:
--
作者:
Steve Kommrusch;Marcos Horro;L. Pouchet;Gabriel Rodríguez;J. Touriño

文献摘要

相似文献

众核处理器具有大量的通用内核,旨在以多线程方式工作。最近的众核处理器使用可扩展的分布式目录保持一致性。一个重要的例子是Intel Mesh互连,它由片上网络互连“瓦片”组成,每个瓦片都包含计算核心、本地缓存和一致性主机。分布式一致性子系统必须针对每个片外访问进行查询,从而对存储器延迟施加开销。本文研究了Intel Knights Landing处理器的物理布局,特别关注一致性子系统,并揭示了分布式目录中物理内存块的伪随机映射功能。利用这些知识,候选优化,以改善内存延迟,通过最小化的一致性流量进行了研究。虽然这些优化确实提高了内存吞吐量,但由于映射函数的计算复杂性所带来的固有开销,最终这并不能转化为性能增益。
Manycore processors feature a high number of general-purpose cores designed to work in a multithreaded fashion. Recent manycore processors are kept coherent using scalable distributed directories. A paramount example is the Intel Mesh interconnect, which consists of a network-on-chip interconnecting “tiles”, each of which contains computation cores, local caches, and coherence masters. The distributed coherence subsystem must be queried for every out-of-tile access, imposing an overhead on memory latency. This paper studies the physical layout of an Intel Knights Landing processor, with a particular focus on the coherence subsystem, and uncovers the pseudo-random mapping function of physical memory blocks across the pieces of the distributed directory. Leveraging this knowledge, candidate optimizations to improve memory latency through the minimization of coherence traffic are studied. Although these optimizations do improve memory throughput, ultimately this does not translate into performance gains due to inherent overheads stemming from the computational complexity of the mapping functions.