Determining Optimal Coherency Interface for Many-Accelerator SoCs Using Bayesian Optimization

Determining Optimal Coherency Interface for Many-Accelerator SoCs Using Bayesian Optimization
复制标题

DOI:
10.1109/lca.2019.2910521
复制
发表时间:
2019-07
影响因子:
2.3
通讯作者:
K. Bhardwaj;Marton Havasi;Yuan Yao;D. Brooks;José Miguel Hernández Lobato;Gu-Yeon Wei
K. Bhardwaj;Marton Havasi;Yuan Yao;D. Brooks;José Miguel Hernández Lobato;Gu-Yeon Wei
中科院分区:
计算机科学3区
文献类型:
--
作者:
K. Bhardwaj;Marton Havasi;Yuan Yao;D. Brooks;José Miguel Hernández Lobato;Gu-Yeon Wei

文献摘要

相似文献

当前的Exascale计算时代的现代芯片(SOC)是复杂的。内存层次结构:非连贯性,与最后级别的缓存(LLC)相干,并且完全连接。 ,加速器可以直接访问主内存,这可能会考虑性能惩罚;而在LLC-Coherent模型中,加速器可以访问LLC,但由于几个加速器之间的争议,可能会遭受性能瓶颈;在私人缓存中,鉴于每个接口的局限为了优化整体系统性能,还提出了目标应用程序。性能。用于图像处理和分类工作负载,与其他“同质”接口相比,混合界面的性能要高出23%,在该界面中,所有加速器都使用单个相干模型。
The modern system-on-chip (SoC) of the current exascale computing era is complex. These SoCs not only consist of several general-purpose processing cores but also integrate many specialized hardware accelerators. Three common coherency interfaces are used to integrate the accelerators with the memory hierarchy: non-coherent,coherent with the last-level cache (LLC), and fully-coherent.However, using a single coherence interface for all the accelerators in an SoC can lead to significant overheads: in the non-coherent model, accelerators directly access the main memory, which can have considerable performance penalty; whereas in the LLC-coherent model, the accelerators access the LLC but may suffer from performance bottleneck due to contention between several accelerators; and the fully-coherent model, that relies on private caches, can incur non-trivial power/area overheads. Given the limitations of each of these interfaces, this paper proposes a novel performance-aware hybrid coherency interface, where different accelerators use different coherency models, decided at design time based on the target applications so as to optimize the overall system performance. A new Bayesian optimization based framework is also proposed to determine the optimal hybrid coherency interface, i.e., use machine learning to select the best coherency model for each of the accelerators in the SoC in terms of performance. For image processing and classification workloads, the proposed framework determined that a hybrid interface achieves up to 23 percent better performance compared to the other ’homogeneous‘ interfaces, where all the accelerators use a single coherency model.