Enhancing Job Scheduling on Inter-Rackscale Datacenters with Free-Space Optical Links

Enhancing Job Scheduling on Inter-Rackscale Datacenters with Free-Space Optical Links
复制标题

DOI:
10.1587/transinf.2018pap0010
复制
发表时间:
2018-12-01
影响因子:
0.7
通讯作者:
Koibuchi, Michihiro
Koibuchi, Michihiro
中科院分区:
计算机科学4区
文献类型:
--
作者:
Hu, Yao;Koibuchi, Michihiro

文献摘要

被引文献

相似文献

数据中心流量和规模的增长正在推动为不同特定应用构建具有低延迟通信的紧密耦合设施的创新。一个著名的定制设计是机架规模(RS)计算,它将关键的服务器资源组件收集到不同的资源池中。这样的资源池实现需要一个新的软件栈来管理资源发现、资源分配和数据通信。为了支持RS中的上述需求,可能需要在其组件上重新配置互连网络。在此背景下,作为原始RS架构的演变,提出了机架间规模(IRS)架构,该架构根据硬件组件所在的区域将其分解到不同的机架中。IRS的核心是使用有限数量的自由空间光学(FSO)通道用于不同资源机架之间的无线连接,通过这些通道,选定的机架对可以直接通信,从而满足资源池需求,而无需额外的软件管理。在本研究中,我们评估了FSO链路对IRS网络的影响。评估结果表明,FSO链路减少了用户作业的平均通信跳数,这接近于2跳的最佳可能值,从而提供了与对应RS架构相当的基准性能。另外,在每个机架配置4个FSO终端的情况下,CPU/SSD (GPU)对接时延比Fat-tree降低25.99%,比2d Torus降低67.14%。我们还介绍了配备fso的IRS系统在给定基准工作负载集的分派作业的平均周转时间方面的优势。
Datacenter growth in traffic and scale is driving innovations in constructing tightly-coupled facilities with low-latency communication for different specific applications. A famous custom design is rackscale (RS) computing by gathering key server resource components into different resource pools. Such a resource-pooling implementation requires a new software stack to manage resource discovery, resource allocation and data communication. The reconfiguration of interconnection networks on their components is potentially needed to support the above demand in RS. In this context as an evolution of the original RS architecture the inter-rackscale (IRS) architecture, which disaggregates hardware components into different racks according to their own areas, has been proposed. The heart of IRS is to use a limited number of free-space optics (FSO) channels for wireless connections between different resource racks, via which selected pairs of racks can communicate directly and thus resource-pooling requirements are met without additional software management. In this study we evaluate the influences of FSO links on IRS networks. Evaluation results show that FSO links reduce average communication hop count for user jobs, which is close to the best possible value of 2 hops and thus provides comparable benchmark performance to that of the counterpart RS architecture. In addition, if four FSO terminals per rack are allowed, the CPU/SSD (GPU) interconnection latency is reduced by 25.99% over Fat-tree and by 67.14% over 2-D Torus. We also present the advantage of an FSO-equipped IRS system in average turnaround time of dispatched jobs for given sets of benchmark workloads.