In-Depth Analysis on Microarchitectures of Modern Heterogeneous CPU-FPGA Platforms

In-Depth Analysis on Microarchitectures of Modern Heterogeneous CPU-FPGA Platforms
复制标题

现代异构CPU-FPGA平台微架构深入分析

DOI:
10.1145/3294054
复制
发表时间:
2019
影响因子:
2.3
通讯作者:
Wei, Peng
Wei, Peng
中科院分区:
计算机科学3区
文献类型:
--
作者:
Choi, Young-Kyu;Cong, Jason;Fang, Zhenman;Hao, Yuchen;Reinman, Glenn;Wei, Peng

文献摘要

参考文献

被引文献

相似文献

传统的同构多核处理器无法提供我们从过去的努力中所期望的持续的性能和能量改进。具有专用硬件加速器的异构架构被广泛认为是解决这个问题的有前途的范例。在不同的异构设备中,可以重新配置以加速具有数量级性能/功耗增益的广泛应用的FPGA正在吸引学术界和工业界越来越多的关注。因此,行业供应商提供了各种具有多样化微架构功能的CPU-FPGA加速平台。然而,这种多样性给应用程序开发人员在为特定应用程序或应用程序域选择合适的平台时带来了严峻的挑战。本文旨在通过确定哪些微体系结构特征会影响性能,以及以何种方式影响性能来解决这一挑战。具体来说,我们对五种最先进的CPU-FPGA加速平台进行了定量比较和深入分析:(1)Alpha Data板和(2)Amazon F1实例,它们代表了传统的基于私有设备内存的PCIe平台;(3)IBM CAPI,代表了基于一致共享内存的PCIe系统;(4)第一代英特尔至强+FPGA加速器平台,代表具有一致共享内存的基于QPI的系统;以及(5)第二代英特尔至强+FPGA加速器平台,代表具有共享内存的基于PCIe(非一致)和基于QPI(一致)的混合系统。通过对CPU与FPGA通信延迟和带宽特性的分析,为应用开发者和平台设计者提供了一系列的见解。此外,我们进行了两个案例研究,以展示如何利用这些见解来优化加速器设计。用于评价的微观基准已公布供公众使用。
Conventional homogeneous multicore processors are not able to provide the continued performance and energy improvement that we have expected from past endeavors. Heterogeneous architectures that feature specialized hardware accelerators are widely considered a promising paradigm for resolving this issue. Among different heterogeneous devices, FPGAs that can be reconfigured to accelerate a broad class of applications with orders-of-magnitude performance/watt gains, are attracting increased attention from both academia and industry. As a consequence, a variety of CPU-FPGA acceleration platforms with diversified microarchitectural features have been supplied by industry vendors. Such diversity, however, poses a serious challenge to application developers in selecting the appropriate platform for a specific application or application domain.This article aims to address this challenge by determining which microarchitectural characteristics affect performance, and in what ways. Specifically, we conduct a quantitative comparison and an in-depth analysis on five state-of-the-art CPU-FPGA acceleration platforms: (1) the Alpha Data board and (2) the Amazon F1 instance that represent the traditional PCIe-based platform with private device memory; (3) the IBM CAPI that represents the PCIe-based system with coherent shared memory; (4) the first generation of the Intel Xeon+FPGA Accelerator Platform that represents the QPI-based system with coherent shared memory; and (5) the second generation of the Intel Xeon+FPGA Accelerator Platform that represents a hybrid PCIe-based (non-coherent) and QPI-based (coherent) system with shared memory. Based on the analysis of their CPU-FPGA communication latency and bandwidth characteristics, we provide a series of insights for both application developers and platform designers. Furthermore, we conduct two case studies to demonstrate how these insights can be leveraged to optimize accelerator designs. The microbenchmarks used for evaluation have been released for public use.
DOI: 10.1145/2678373.2665678
发表时间: 2014-10
期刊: 2014 ACM/IEEE 41st International Symposium on Computer Architecture (ISCA)
影响因子: --
作者:
Andrew Putnam;Adrian M. Caulfield;Eric S. Chung;Derek Chiou;Kypros Constantinides;J. Demme;H. Esmaeilzadeh;J. Fowers;Gopi Prashanth Gopal;J. Gray;M. Haselman;S. Hauck;Stephen Heil;Amir Hormati;Joo-Young Kim;S. Lanka;J. Larus;Eric Peterson;Simon Pope;Aaron Smith;J. Thong;Phillip Yi Xiao;D. Burger
通讯作者: Andrew Putnam;Adrian M. Caulfield;Eric S. Chung;Derek Chiou;Kypros Constantinides;J. Demme;H. Esmaeilzadeh;J. Fowers;Gopi Prashanth Gopal;J. Gray;M. Haselman;S. Hauck;Stephen Heil;Amir Hormati;Joo-Young Kim;S. Lanka;J. Larus;Eric Peterson;Simon Pope;Aaron Smith;J. Thong;Phillip Yi Xiao;D. Burger
基于缓存一致性结构的可重构计算系统
DOI: 10.1109/reconfig.2011.4
发表时间: 2011
期刊: 2011 International Conference on Reconfigurable Computing and FPGAs
影响因子: --
作者:
Neal Oliver;Rahul R. Sharma;Stephen Chang;Bhushan Chitlur;E. Garcia;Joseph Grecco;Aaron Grier;Nelson Ijih;Yaping Liu;Pratik Marolia;H. Mitchel;S. Subhaschandra;Arthur Sheiman;Timothy S. Whisonant;Prabhat Gupta
通讯作者: Prabhat Gupta
DOI: 10.1109/fccm.2018.00011
发表时间: 2018-04
期刊: 2018 IEEE 26th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)
影响因子: --
作者:
Zhenyuan Ruan;Tong He;Bojie Li;Peipei Zhou;J. Cong
通讯作者: Zhenyuan Ruan;Tong He;Bojie Li;Peipei Zhou;J. Cong
DOI: 10.1147/jrd.2014.2380198
发表时间: 2015-02
期刊: IBM J. Res. Dev.
影响因子: --
作者:
Jeffrey Stuecheli;B. Blaner;C. Johns;M. S. Siegel
通讯作者: Jeffrey Stuecheli;B. Blaner;C. Johns;M. S. Siegel
从 JVM 到 FPGA:通过优化的深度流水线桥接抽象层次结构
DOI: --
发表时间: 2018
期刊: The 10th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 2018
影响因子: --
作者:
Cong, Jason;Wei, Peng;Yu, Cody Hao
通讯作者: Yu, Cody Hao