CAREER: Architecting Datacenters for Optimized Tail Latency at Scale
CAREER: Architecting Datacenters for Optimized Tail Latency at Scale
批准号:
2237434
负责人:
Alexandros Daglis
金额:
$53.55万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2028-09-30
中文摘要
现代社会严重依赖云计算和各种在线服务,例如社交媒体、网络搜索、数字内容交付等。这些实用程序的数十亿美元市场是通过大规模部署服务器(即,企业级计算机),称为数据中心。对于这些企业部署的服务来说,必须对每个用户请求可靠地生成高质量、低延迟的响应,因为即使是偶尔的性能或可用性故障也会导致重大的收入损失。因此,服务提供者设置了严格的服务级别目标(SLO),用于定义每个服务组件的响应延迟分布的末端的可接受行为。遵守这样的SLO是具有挑战性的,因为计算系统通常被优化以满足平均而不是尾部性能目标。第二个关键挑战来自于需要跟上下一代服务不断增长的需求。服务器不断通过数据中心的内部网络进行通信,以在功能、用户和数据集方面以所需的规模协作地启用在线服务。随着这三个方面的需求不断增长,据报道,数据中心内的服务器间数据移动每12-15个月翻一番,因此成为一个主要的性能决定因素。该项目旨在为未来网络中心的设计提供信息,这些网络中心将能够跟上为现代数字经济提供动力的下一代在线服务日益增长的需求。实现这一目标的两种主要方法是(i)通过整体跨组件来大幅提高数据中心内的数据移动效率(计算、内存和网络)设计优化,(二)发展中国家--嵌入在每个关键系统组件中的感知机制,以原生地满足数据中心环境特有的感兴趣的性能指标。该项目追求两个主要途径来提高下一代互联网服务提供商在提供高质量在线服务方面发挥着关键作用。第一条途径调查的潜力,促进SLO意识从一个回顾性的评估指标,以普遍的,整体优化旋钮从集群规模下降到微架构。将开发一个整体的SLO感知优化框架,以实现集群范围的动态请求优先级策略的执行。反过来,这些策略将驱动在性能关键系统组件(包括计算、网络和内存资源)的底层硬件中实现的一组SLO感知机制。由于数据移动主要决定计算效率,第二种研究途径提出了在集群规模和每个单独端点(即,服务器)。随着网络功能的增长,服务器上的数据移动必须通过每个服务器的网络接口和内存层次结构的明智的协同设计来智能地协调。这两种研究途径相结合,引入了一个新的大规模系统设计的角度来看,一个整体的系统设计方法,使新的组件间的协同作用,承诺大幅提高端到端的性能和效率。最后,计划研究的第三个目标,两个主要研究途径的组成部分,是开发新的仿真方法和工具,使企业级技术的评估使用有限的计算资源在典型的学术环境中实现,同时达到模拟系统的规模,速度和准确性之间的平衡。该奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Modern societies heavily rely on cloud computing and a wide variety of online services, such as social media, web search, digital content delivery, etc. The multi-billion-dollar market of these utilities is enabled by massive deployments of servers (i.e., enterprise-grade computers), known as datacenters. It is imperative for such datacenter-deployed services to reliably generate high-quality, low-latency responses to every user request, as even infrequent performance or availability hiccups incur significant revenue loss. Therefore, service providers set strict Service Level Objectives (SLOs) that define the acceptable behavior of the tail-end of each service component's response latency distribution. Abiding by such SLOs is challenging, as computing systems have been conventionally optimized to meet average rather than tail performance goals. A second key challenge stems from the need to keep up with growing demands of next-generation services. Servers constantly communicate over the datacenter's internal network to collaboratively enable an online service at the required scale, in terms of features, users, and datasets. With demand relentlessly growing in each of these three dimensions, inter-server data movement within the datacenter is reportedly doubling every 12-15 months, thus becoming a major performance determinant. This project aims to inform the design of future datacenters that will be able to keep up with the growing demands of next-generation online services powering modern digital economies. The two primary approaches to achieve that are (i) drastic improvement of intra-datacenter data movement efficiency via holistic cross-component (compute, memory, and network) design optimizations, and (ii) development of SLO-aware mechanisms embedded within each key system component to natively cater to the performance metrics of interest that are unique to datacenter environments.This project pursues two main avenues to advance the efficiency of next-generation datacenters in their crucial role of delivering high-quality online services. The first avenue investigates the potential of promoting SLO-awareness from a retrospective evaluation metric to a pervasive, integral optimization knob from cluster scale down to microarchitecture. A holistic SLO-aware optimization framework will be developed to enable enforcement of cluster-wide dynamic request prioritization policies. In turn, these policies will drive a set of SLO-aware mechanisms implemented in the underlying hardware of performance-critical system components, including compute, network, and memory resources. Because data movement predominantly dictates computation efficiency, the second research avenue proposes techniques for reduced data movement both at the cluster scale and at each individual endpoint (i.e., server). As networking capabilities grow, on-server data movement must be intelligently orchestrated via judicious co-design of each server's network interface and memory hierarchy. The two research avenues combined introduce a new large-scale system design perspective, where a holistic system design approach enables new inter-component synergies that promise drastic improvements to end-to-end performance and efficiency. Finally, a third objective of the planned research, integral to the two main research avenues, is the development of new simulation methodologies and tools that enable evaluation of datacenter-scale techniques using limited compute resources attainable in typical academic environments, while striking a balance between simulated system scale, speed, and accuracy. Synergistic datacenter-themed educational activities will also be undertaken at the graduate, undergraduate, and K-12 levels.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SHF: Small: Redesigning the Memory System in the Era of Compute Express Link
-
批准号:2333049
-
项目类别:Standard Grant
-
资助金额:$57.0万
-
财政年份:2024
-
负责人:Alexandros Daglis
-
依托单位:
SHF: CNS Core: Small: Server architecture optimizations for microsecond-scale RPCs
-
批准号:2006602
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2020
-
负责人:Alexandros Daglis
-
依托单位:
海外基金