Manycore Network Interfaces for in-memory rack-scale computing

Manycore Network Interfaces for in-memory rack-scale computing
复制标题

用于内存机架规模计算的众核网络接口

DOI:
10.1145/2749469.2750415
复制
发表时间:
2015
期刊:
2015 ACM/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Boris Grot
Boris Grot
中科院分区:
--
文献类型:
--
作者:
Alexandros Daglis;Stanko Novakovic;Edouard Bugnion;B. Falsafi;Boris Grot

文献摘要

被引文献

相似文献

数据中心操作员依靠低成本,高密度技术来最大程度地提高吞吐量,以使用紧密的尾部潜伏期进行数据密集型服务。内存中的机架规模计算正在成为大规模数据中心的有前途的范式,该数据中心资本利用商品SOC,低延迟和高带宽通信面料和远程存储器访问模型,以促进机架内存,以使机架的内存用于关键数据密集型的存储器诸如图形处理或键值存储之类的应用程序。低潜伏期和高带宽不仅决定了软件协议和片外织物中消除通信瓶颈的决定,而且还决定了网络界面的芯片芯片集成。后者是一个关键挑战,尤其是在具有RDMA启发的单方面操作的体系结构中,旨在通过芯片网络接口(NI)支持实现低潜伏期和高带宽。本文提出并评估了用于内存机架规模计算的瓷砖多方SOC的网络接口体系结构。我们的结果表明,每个芯片瓷砖的Ni功能仔细分裂,沿NOC尺寸在芯片的边缘处逐渐分配,使机架规模的体系结构可以优化延迟和带宽。我们最好的Many-Core NI体系结构可在理想化的硬件NUMA的3%内达到潜伏期,并有效地使用NOC的完整分配带宽,而无需更改芯片连贯协议或Core的微结构。
Datacenter operators rely on low-cost, high-density technologies to maximize throughput for data-intensive services with tight tail latencies. In-memory rack-scale computing is emerging as a promising paradigm in scale-out datacenters capitalizing on commodity SoCs, low-latency and high-bandwidth communication fabrics and a remote memory access model to enable aggregation of a rack's memory for critical data-intensive applications such as graph processing or key-value stores. Low latency and high bandwidth not only dictate eliminating communication bottlenecks in the software protocols and off-chip fabrics but also a careful on-chip integration of network interfaces. The latter is a key challenge especially in architectures with RDMA-inspired one-sided operations that aim to achieve low latency and high bandwidth through on-chip Network Interface (NI) support. This paper proposes and evaluates network interface architectures for tiled manycore SoCs for in-memory rack-scale computing. Our results indicate that a careful splitting of NI functionality per chip tile and at the chip's edge along a NOC dimension enables a rack-scale architecture to optimize for both latency and bandwidth. Our best manycore NI architecture achieves latencies within 3% of an idealized hardware NUMA and efficiently uses the full bisection bandwidth of the NOC, without changing the on-chip coherence protocol or the core's microarchitecture.