Programmable FPGA-based Memory Controller

Programmable FPGA-based Memory Controller
复制标题

DOI:
10.1109/hoti52880.2021.00020
复制
发表时间:
2021-08
期刊:
2021 IEEE Symposium on High-Performance Interconnects (HOTI)
影响因子:
--
通讯作者:
Sasindu Wijeratne;S. Pattnaik;Zhiyu Chen;R. Kannan;V. Prasanna
Sasindu Wijeratne;S. Pattnaik;Zhiyu Chen;R. Kannan;V. Prasanna
中科院分区:
其他
文献类型:
--
作者:
Sasindu Wijeratne;S. Pattnaik;Zhiyu Chen;R. Kannan;V. Prasanna

文献摘要

相似文献

即使在DRAM技术方面有一代改进,内存访问潜伏期仍然是应用程序加速器的主要瓶颈,这主要是由于内存界面IPS的限制无法完全考虑目标应用程序的变化,所使用的算法和加速器架构架构。由于为不同应用程序开发内存控制器是耗时的,因此本文引入了一个模块化和可编程的内存控制器,可以在可用的硬件资源上为不同的目标应用程序配置。提出的内存控制器有效地支持缓存线访问以及批量内存传输。用户可以根据FPGA,内存访问模式和外部内存规范的可用逻辑资源来配置控制器。模块化设计支持各种内存访问优化技术,包括,请求调度,内部缓存和直接内存访问。这些技术有助于减少整体延迟,同时保持高持续的带宽。我们在最先进的FPGA上实现该系统,并使用两个广泛研究的域评估其性能:图形分析和深度学习工作负载。与商业内存控制器IP相比,我们显示了CNN和GCN工作量最高58%的总体内存访问时间。
Even with generational improvements in DRAM technology, memory access latency still remains the major bottleneck for application accelerators, primarily due to limitations in memory interface IPs which cannot fully account for variations in target applications, the algorithms used, and accelerator architectures. Since developing memory controllers for different applications is time-consuming, this paper introduces a modular and programmable memory controller that can be configured for different target applications on available hardware resources. The proposed memory controller efficiently supports cache-line accesses along with bulk memory transfers. The user can configure the controller depending on the available logic resources on the FPGA, memory access pattern, and external memory specifications. The modular design supports various memory access optimization techniques including, request scheduling, internal caching, and direct memory access. These techniques contribute to reducing the overall latency while maintaining high sustained bandwidth. We implement the system on a state-of-the-art FPGA and evaluate its performance using two widely studied domains: graph analytics and deep learning workloads. We show improved overall memory access time up to 58% on CNN and GCN workloads compared with commercial memory controller IPs.