Towards Application-Specific Address Mapping for Emerging Memory Devices

Towards Application-Specific Address Mapping for Emerging Memory Devices
复制标题

DOI:
10.1145/3422575.3422785
复制
发表时间:
2020-09
期刊:
Proceedings of the International Symposium on Memory Systems
影响因子:
--
通讯作者:
Shashank Adavally
Shashank Adavally
中科院分区:
其他
文献类型:
--
作者:
Shashank Adavally

文献摘要

相似文献

3d堆叠DRAM的最新进展,如混合内存立方体(HMC)和高带宽内存(HBM),与传统的基于ddr的DRAM相比,有望实现更高的带宽和更低的功耗。然而,利用这些额外的带宽来提高实际应用程序的性能需要仔细地在内存中布局数据,这需要程序员付出大量的努力。为了减轻程序员的负担,我们研究了特定于应用程序的地址映射,以提高性能,同时最大限度地减少手工工作。我们的方法是由以下见解指导的:(i)地址位的切换活动可以帮助确定提高内存内并行性的策略,但这个指标低估了冲突;(ii)现代内存控制器重新排序地址请求,因此从地址跟踪测量的任何切换活动都是不确定的。此外,我们的立场是,分析单个地址比特会导致对实际冲突和利用并行性的不良估计,并且需要计算地址比特组的熵。因此,我们计算了一组地址位的基于窗口的概率熵,以确定一个接近最优的地址映射。我们给出了十个应用程序的模拟结果,表明我们提出的方法的性能比固定地址映射提高了25%,比以前的特定于应用程序的地址映射提高了8%。
Recent advancements in 3D-stacked DRAM such as hybrid memory cube (HMC) and high-bandwidth memory (HBM) promise higher bandwidth and lower power consumption compared to traditional DDR-based DRAM. However, taking advantage of this additional bandwidth for improving the performance of real-world applications requires carefully laying out the data in memory which incurs significant programmer effort. To alleviate this programmer burden, we investigate application-specific address mapping to improve performance while minimizing manual effort. Our approach is guided by the following insights: (i) toggling activity of address bits can help determine strategies to improve parallelism within memory but this metric underestimates conflicts and (ii) modern memory controllers reorder address requests and therefore any toggling activity measured from an address trace is non-deterministic. Furthermore, our position is that analyzing individual address bits results in poor estimates for actual conflicts and exploited parallelism and that entropy needs to be calculated for groups of address bits. Therefore, we calculate window-based probabilistic entropy for groups of address bits to determine a near-optimal address mapping. We present simulation results for ten applications that show a performance improvement up to 25% over fixed address-mapping and up to 8% over previous application-specific address mapping for our proposed approach.