DynaBurst: Dynamically Assemblying DRAM Bursts over a Multitude of Random Accesses

DynaBurst: Dynamically Assemblying DRAM Bursts over a Multitude of Random Accesses
复制标题

DynaBurst:动态组装 DRAM 在多次随机访问中突发

DOI:
--
复制
发表时间:
2019
期刊:
International Conference on Field-Programmable Logic and Applications
影响因子:
--
通讯作者:
P. Ienne
P. Ienne
中科院分区:
--
文献类型:
--
作者:
Mikhail Asiatici;P. Ienne

文献摘要

被引文献

相似文献

FPGA外部内存的有效带宽(通常是DRAM)对访问模式极为敏感。每当内存访问不规则且不可用,或者在设计时间方面不昂贵时,处理数千个出色的失误(错过的内存系统)的非阻止库可以动态地改善带宽利用率。但是,它们需要一个在FPGA侧具有宽数据端口的内存控制器,并且无法完全利用具有SOC FPGA上常见的多个狭窄端口的内存接口。此外,由于它们的范围仅限于单个内存请求,因此它们生成的访问模式可能会导致频繁的DRAM行冲突,从而进一步减少DRAM带宽。在本文中,我们提出了dynaburst,这是一个错过的内存系统的扩展,该系统将可变的长度爆发向内存控制器生成。通过使内存访问在本地更顺序,我们可以最大程度地减少DRAM行冲突的数量,并以每次要求的基础调整突发长度,我们可以最大程度地减少带宽浪费。在多个狭窄的DDR3控制器上,我们提供28%的几何平均值,并且与同一区域的传统非封锁缓存相比,高达3.4倍的速度,而先前的单要求方法并不具有成本效益。在具有单个宽端口的控制器上,我们可以进一步提高错过的系统的性能,最高可达2.4倍。
The effective bandwidth of the FPGA external memory, usually DRAM, is extremely sensitive to the access pattern. Nonblocking caches that handle thousands of outstanding misses (miss-optimized memory systems) can dynamically improve bandwidth utilization whenever memory accesses are irregular and application-specific optimizations are not available or are too costly in terms of design time. However, they require a memory controller with wide data ports on the FPGA side and cannot fully take advantage of the memory interfaces with multiple narrow ports that are common on SoC FPGAs. Moreover, as their scope is limited to single memory requests, the access pattern they generate may cause frequent DRAM row conflicts, which further reduce DRAM bandwidth. In this paper, we propose DynaBurst, an extension of miss-optimized memory systems that generates variable-length bursts to the memory controller. By making memory accesses locally more sequential, we minimize the number of DRAM row conflicts, and by adapting the burst length on a per-request basis we minimize bandwidth wastage. On a multiple, narrow-ported DDR3 controller, we provide 28% geometric mean and up to 3.4x speedup compared to a traditional nonblocking cache of the same area, while the prior single-request approach would not have been cost-effective. On a controller with a single, wide port, we can further improve the performance of miss-optimized systems by up to 2.4x.