LPM: Concurrency-Driven Layered Performance Matching

LPM: Concurrency-Driven Layered Performance Matching
复制标题

DOI:
10.1109/icpp.2015.97
复制
发表时间:
2015-09
期刊:
2015 44th International Conference on Parallel Processing
影响因子:
--
通讯作者:
Yuhang Liu;Xian-He Sun
Yuhang Liu;Xian-He Sun
中科院分区:
其他
文献类型:
--
作者:
Yuhang Liu;Xian-He Sun

文献摘要

被引文献

相似文献

数据访问已成为影响计算性能的突出瓶颈。本文提出了一种分层性能匹配(LPM)模型及其相关算法,以匹配存储层次结构中每一层的请求和响应速度,从而提高存储性能。LPM的基本原理是,内存层次结构的每一层的性能应该而且可以进行优化,以与其正上方的层的请求紧密匹配。LPM模型同时考虑了数据访问的并发性和局部性。它揭示了这样一个事实:增加较高层的命中和未命中之间的有效重叠将缓解较低层的性能影响。引入了纯未命中和纯未命中惩罚两个术语来衡量这种命中-未命中重叠的有效性。通过区分(一般)脱靶量和纯脱靶量,使LPM优化具有实用性和可行性。我们的评估表明,通过优化硬件配置,可以显著减少数据停顿时间。我们还通过简单地采用智能LPM调度而无需更改底层硬件配置,实现了显著的性能提升。分析和实验结果表明,LPM是可行和有效的。它为解决日益扩大的存储墙问题,优化至关重要的存储系统设计提供了一种新颖而有效的方法。
Data access has become the preeminent performance bottleneck of computing. In this study, a Layered Performance Matching (LPM) model and its associated algorithm are proposed to match the request and reply speed for each layer of a memory hierarchy to improve memory performance. The rationale of LPM is that the performance of each layer of a memory hierarchy should and can be optimized to closely match the request of the layer directly above it. The LPM model simultaneously considers both data access concurrency and locality. It reveals the fact that increasing the effective overlapping between hits and misses of the higher layer will alleviate the performance impact of the lower layer. The terms pure miss and pure miss penalty are introduced to measure the effectiveness of such hit-miss overlapping. By distinguishing between (general) miss and pure miss, we have made LPM optimization practical and feasible. Our evaluation shows the data stall time can be reduced significantly with an optimized hardware configuration. We also have achieved noticeable performance improvement by simply adopting smart LPM scheduling without changing the underlying hardware configurations. Analysis and experimental results show LPM is feasible and effective. It provides a novel and efficient way to cope with the ever-widening memory wall problem, and to optimize the vital memory system design.