A reconfigurable fabric for accelerating large-scale datacenter services

A reconfigurable fabric for accelerating large-scale datacenter services
复制标题

DOI:
10.1145/2678373.2665678
复制
发表时间:
2014-10
期刊:
2014 ACM/IEEE 41st International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Andrew Putnam;Adrian M. Caulfield;Eric S. Chung;Derek Chiou;Kypros Constantinides;J. Demme;H. Esmaeilzadeh;J. Fowers;Gopi Prashanth Gopal;J. Gray;M. Haselman;S. Hauck;Stephen Heil;Amir Hormati;Joo-Young Kim;S. Lanka;J. Larus;Eric Peterson;Simon Pope;Aaron Smith;J. Thong;Phillip Yi Xiao;D. Burger
Andrew Putnam;Adrian M. Caulfield;Eric S. Chung;Derek Chiou;Kypros Constantinides;J. Demme;H. Esmaeilzadeh;J. Fowers;Gopi Prashanth Gopal;J. Gray;M. Haselman;S. Hauck;Stephen Heil;Amir Hormati;Joo-Young Kim;S. Lanka;J. Larus;Eric Peterson;Simon Pope;Aaron Smith;J. Thong;Phillip Yi Xiao;D. Burger
中科院分区:
其他
文献类型:
--
作者:
Andrew Putnam;Adrian M. Caulfield;Eric S. Chung;Derek Chiou;Kypros Constantinides;J. Demme;H. Esmaeilzadeh;J. Fowers;Gopi Prashanth Gopal;J. Gray;M. Haselman;S. Hauck;Stephen Heil;Amir Hormati;Joo-Young Kim;S. Lanka;J. Larus;Eric Peterson;Simon Pope;Aaron Smith;J. Thong;Phillip Yi Xiao;D. Burger

文献摘要

被引文献

相似文献

数据中心工作负载需要高计算能力、灵活性、能效和低成本。同时改善所有这些因素是具有挑战性的。为了提升数据中心的功能,使其超出商品服务器设计所能提供的范围,我们设计并构建了一个可组合、可重新配置的结构,以加速部分大规模软件服务。该结构的每个实例都由嵌入到48台机器的半机架中的6×8 2-D高端Stratix V FPGA圆环组成。每个服务器中放置一个FPGA,可通过PCIe访问,并通过成对的10 Gb SAS电缆直接连接到其他FPGA。在本文中,我们描述了在1,632台服务器的床上部署此结构的中等规模部署,并测量了其在加速Bing Web搜索引擎方面的功效。我们描述了系统的要求和体系结构,详细说明了关键的工程挑战和解决方案,使系统在出现故障时保持稳健,并在对候选文档进行排名时测量系统的性能,功率和弹性。在高负载下,对于固定的延迟分布,大规模可重配置结构将每个服务器的排名吞吐量提高了95%,或者在保持相等吞吐量的同时,将尾部延迟降低了29%。
Datacenter workloads demand high computational capabilities, flexibility, power efficiency, and low cost. It is challenging to improve all of these factors simultaneously. To advance datacenter capabilities beyond what commodity server designs can provide, we have designed and built a composable, reconfigurable fabric to accelerate portions of large-scale software services. Each instantiation of the fabric consists of a 6×8 2-D torus of high-end Stratix V FPGAs embedded into a half-rack of 48 machines. One FPGA is placed into each server, accessible through PCIe, and wired directly to other FPGAs with pairs of 10 Gb SAS cables. In this paper, we describe a medium-scale deployment of this fabric on a bed of 1,632 servers, and measure its efficacy in accelerating the Bing web search engine. We describe the requirements and architecture of the system, detail the critical engineering challenges and solutions needed to make the system robust in the presence of failures, and measure the performance, power, and resilience of the system when ranking candidate documents. Under high load, the largescale reconfigurable fabric improves the ranking throughput of each server by a factor of 95% for a fixed latency distribution-or, while maintaining equivalent throughput, reduces the tail latency by 29%.