A comprehensive methodology to determine optimal coherence interfaces for many-accelerator SoCs

A comprehensive methodology to determine optimal coherence interfaces for many-accelerator SoCs
复制标题

DOI:
10.1145/3370748.3406564
复制
发表时间:
2020-08
期刊:
Proceedings of the ACM/IEEE International Symposium on Low Power Electronics and Design
影响因子:
--
通讯作者:
K. Bhardwaj;Marton Havasi;Yuan Yao;D. Brooks;José Miguel Hernández-Lobato;Gu-Yeon Wei
K. Bhardwaj;Marton Havasi;Yuan Yao;D. Brooks;José Miguel Hernández-Lobato;Gu-Yeon Wei
中科院分区:
其他
文献类型:
--
作者:
K. Bhardwaj;Marton Havasi;Yuan Yao;D. Brooks;José Miguel Hernández-Lobato;Gu-Yeon Wei

文献摘要

相似文献

现代片上系统(SOCS)不仅包括通用CPU,还包括专门的硬件加速器。通常,有三种连贯的模型选择可以将加速器与内存层次结构集成:没有连贯性,与最后级别的缓存(LLC)和基​​于私有缓存的完整相干性相干。但是,发现哪种相干模型对于复杂的多加速器SOC的加速器的最佳研究非常有限。本文着重于确定SOC及其目标应用的成本感知相干界面:为加速器找到优化其功率和性能的最佳连贯模型,考虑到工作负载特征和系统级别的争论。提出了一种新颖的综合方法,该方法使用贝叶斯优化来有效地找到使用GEM5-Aladdin架构模拟器建模的SOC的成本感知相干接口。为了进行完整的分析,除了已经支持的没有连贯性和完全连贯性外,GEM5-Aladdin还扩展了LLC相干性。对于具有不同量的加速器级并行性的异质SOC靶向应用,拟议的框架迅速找到了成本感知的连贯界面,这些框架表现出对其他常用的相干界面显示出明显的性能和功率收益。
Modern systems-on-chip (SoCs) include not only general-purpose CPUs but also specialized hardware accelerators. Typically, there are three coherence model choices to integrate an accelerator with the memory hierarchy: no coherence, coherent with the last-level cache (LLC), and private cache based full coherence. However, there has been very limited research on finding which coherence models are optimal for the accelerators of a complex many-accelerator SoC. This paper focuses on determining a cost-aware coherence interface for an SoC and its target application: find the best coherence models for the accelerators that optimize their power and performance, considering both workload characteristics and system-level contention. A novel comprehensive methodology is proposed that uses Bayesian optimization to efficiently find the cost-aware coherence interfaces for SoCs that are modeled using the gem5-Aladdin architectural simulator. For a complete analysis, gem5-Aladdin is extended to support LLC coherence in addition to already-supported no coherence and full coherence. For a heterogeneous SoC targeting applications with varying amount of accelerator-level parallelism, the proposed framework rapidly finds cost-aware coherence interfaces that show significant performance and power benefits over the other commonly-used coherence interfaces.