SHF: CNS Core: Small: Server architecture optimizations for microsecond-scale RPCs
SHF: CNS Core: Small: Server architecture optimizations for microsecond-scale RPCs
批准号:
2006602
负责人:
Alexandros Daglis
金额:
$40.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2024-09-30
中文摘要
现代数据中心托管在线服务,这些服务被分解为跨越数千台服务器的多个软件层。服务器之间使用内部数据中心网络上的远程过程调用(RPC)进行协调和通信。微服务正在不断提高生产力的软件体系结构趋势正在将已部署服务的软件分解推向极端,导致服务器间通信更加频繁,RPC更加细粒度,运行时间通常只有几微秒。随着每个RPC运行时间的缩减,网络效率直接影响在线服务的整体性能:与网络相关的开销本来可以忽略不计,但每个RPC触发的应用程序级逻辑的细粒度特性放大了这些开销。解决这一挑战的一个有希望的方法是提升每个服务器的NIC-服务器计算资源和网络之间的网关-的作用,从简单的RPC交付到主动RPC操作。过去,NIC通过将传入的数据包写入内存来以不可知性的方式传递它们;这些数据包稍后由CPU核心拾取进行处理,从而导致过度的数据移动、核心间同步开销或核心间负载失衡。Roar是一种新的服务器架构,针对微秒级RPC的高效处理进行了优化,具有一个NIC,可以动态监控系统范围的行为,并智能地引导服务器内存层次结构中的传入RPC,包括直接放置在CPU核心的专用缓存中。面向RPC的协议允许NIC将其操作的抽象级别从网络分组提升到RPC,即从数据块到触发某些计算的消息。向NIC公开的有关RPC将触发的计算的信息越多,NIC可以做出的RPC控制决策就越好。在Roar体系结构中,NIC监视传入的RPC并做出许多新的决策,以便明智地将它们分布在现代服务器的内存层次结构中和跨CPU核心。总体而言,Roar的技术可以极大地提高处理微秒级RPC的效率和性能。一个直接的结果是大量部署在现代数据中心的大规模在线服务的质量提高,这些数据中心大量使用这种RPC。因此,ROAR有可能影响未来服务器架构的设计。ROAR涉及广泛的软硬件协同设计,打破了传统上网络和计算之间僵硬的边界。NIC的角色从不经意间将传入的RPC放置到服务器的内存层次结构中提升为主动RPC加速。NIC在RPC处理生命周期中的自然位置使其成为优化数据移动和降低延迟的缓存层次结构的优秀代理。Roar有三个主要机制。首先,它通过将所有传入RPC保持在缓存层次结构中,提前拒绝由于过多的持续排队而预计将错过截止日期的请求,从而缓解了有害的内存带宽干扰。其次,Roar通过考虑实时系统负载信息,为传入的RPC做出动态的核心间平衡决策,并将RPC一直引导到所选CPU核心的私有缓存。第三,当RPC排队等待执行时,NIC预取RPC的相应指令和关键数据,从而在RPC最终被内核获取以进行处理时加快RPC的启动时间。这种预取的性质是新颖的,因为决策不是基于预测,而是基于先见之明:NIC对RPC从网络到达的早期了解。建议的研究包括理论建模、仿真和原型制作。将对各种RPC服务时间分布进行排队分析,以制定NIC驱动的核心间负载分配策略。为了评估缓存中的网络缓冲区管理、RPC-to-core转向和预取机制,将开发一个周期精确的Roar仿真模型。最后,将使用可编程的基于FPGA的NIC来评估NIC驱动的负载平衡策略在具有离散NIC的现有体系结构上的适用性。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Modern datacenters host online services that are decomposed into multiple software tiers spanning thousands of servers. Servers coordinate and communicate with each other using Remote Procedure Calls (RPCs) over the internal datacenter network. The ongoing productivity-boosting software architecture trend of microservices is pushing software decomposition of deployed services to the extreme, resulting in more frequent inter-server communication and finer-grained RPCs, often with runtimes of only a few microseconds. With shrinking per-RPC runtime, networking efficiency directly impacts the performance of an online service as a whole: networking-related overheads that would otherwise be negligible are amplified by the fine-grained nature of the application-level logic triggered per RPC. A promising approach to address this challenge is to promote the role of each server’s NIC—the gateway between a server’s compute resources and the network—from simple RPC delivery to active RPC manipulation. Historically, the NIC agnostically delivers incoming packets, by writing them into memory; the packets are later picked up by a CPU core for processing, resulting in excess data movement, inter-core synchronization overheads, or inter-core load imbalance. ROAr is a new server architecture optimized for efficient handling of microsecond-scale RPCs, featuring a NIC that dynamically monitors system-wide behavior and intelligently steers incoming RPCs within the server’s memory hierarchy, including direct placement in CPU cores’ private caches. An RPC-oriented protocol allows the NIC to raise the level of abstraction it operates on from network packets to RPCs—i.e., from data chunks to messages that trigger some computation. The more information exposed to the NIC about the computation an RPC will trigger, the better the RPC steering decision the NIC can make. In the ROAr architecture, the NIC monitors incoming RPCs and makes a number of novel decisions to judiciously distribute them within a modern server’s memory hierarchy and across CPU cores. Overall, ROAr’s techniques can drastically improve the efficiency and performance of handling microsecond-scale RPCs. A direct consequence is improved quality for a plethora of large-scale online services deployed on modern datacenters, which make heavy use of such RPCs. Therefore, ROAr has the potential to influence the design of future server architectures.ROAr involves extensive hardware-software co-design, breaking the conventionally rigid boundaries between network and compute. The NIC’s role is promoted from oblivious placement of incoming RPCs into a server’s memory hierarchy to active RPC acceleration. The NIC’s natural position in an RPC’s processing lifetime establishes it as an excellent agent to stage the cache hierarchy for optimized data movement and reduced latency. ROAr features three main mechanisms. First, it alleviates detrimental memory-bandwidth interference by keeping all incoming RPCs within the cache hierarchy, early-rejecting requests that are predicted to miss their deadline because of excessive ongoing queuing. Second, ROAr makes dynamic inter-core balancing decisions for incoming RPCs, by taking into account real-time system load information, and steers RPCs all the way to the private cache of the selected CPU core. Third, while an RPC is queued, waiting to be executed, the NIC prefetches the RPC’s corresponding instructions and critical data, thus accelerating the RPC’s startup time when it is eventually picked up by the core for processing. The nature of such prefetching is novel, as decisions are not based on predictions, but on prescience: the NIC’s early knowledge of an RPC’s arrival from the network. The proposed research involves theoretical modeling, simulation, and prototyping. Queuing analysis on a variety of RPC service-time distributions will be conducted to develop NIC-driven inter-core load distribution policies. A cycle-accurate simulation model of ROAr will be developed to evaluate in-cache network buffer management, RPC-to-core steering, and prefetching mechanisms. Finally, the applicability of NIC-driven load-balancing policies on existing architectures featuring discrete NICs will be evaluated, with the use of a programmable FPGA-based NIC.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1145/3567955.3567957
发表时间:
2022-12
期刊:
Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1
影响因子:
--
作者:
[Mark Sutherland;B. Falsafi;Alexandros Daglis]
通讯作者:
Mark Sutherland;B. Falsafi;Alexandros Daglis
Cerebros: Evading the RPC Tax in Datacenters
Cerebros:逃避数据中心的 RPC 税
DOI:
10.1145/3466752.3480055
发表时间:
2021
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
作者:
[Pourhabibi, Arash, Sutherland, Mark, Daglis, Alexandros, Falsafi, Babak]
通讯作者:
Falsafi, Babak
IDIO: Orchestrating Inbound Network Data on Server Processors
IDIO:在服务器处理器上编排入站网络数据
DOI:
10.1109/lca.2020.3044923
发表时间:
2021
期刊:
IEEE Computer Architecture Letters
影响因子:
2.3
作者:
[Alian, Mohammad, Shin, Jongmin, Kang, Ki-Dong, Wang, Ren, Daglis, Alexandros, Kim, Daehoon, Kim, Nam Sung]
通讯作者:
Kim, Nam Sung
DOI:
10.1109/micro56248.2022.00041
发表时间:
2022-10
期刊:
2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
作者:
[Marina Vemmou;Albert Cho;Alexandros Daglis]
通讯作者:
Marina Vemmou;Albert Cho;Alexandros Daglis
SHF: Small: Redesigning the Memory System in the Era of Compute Express Link
-
批准号:2333049
-
项目类别:Standard Grant
-
资助金额:$57.0万
-
财政年份:2024
-
负责人:Alexandros Daglis
-
依托单位:
CAREER: Architecting Datacenters for Optimized Tail Latency at Scale
-
批准号:2237434
-
项目类别:Continuing Grant
-
资助金额:$53.55万
-
财政年份:2023
-
负责人:Alexandros Daglis
-
依托单位:
国内基金
海外基金
登录
查看更多内容
IL-17A通过STAT5影响CNS2区域甲基化抑制调节性T细胞功能在银屑病发病中的作用和机制研究
-
批准号:82304006
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:刘阳禾
-
依托单位:
miR-20a通过调控CD4+T细胞焦亡促进CNS炎性脱髓鞘疾病的发生及机制研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:王亦舒
-
依托单位:
血浆CNS来源外泌体中寡聚磷酸化α-synuclein对PD病程的提示研究
-
批准号:82101506
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:徐妍
-
依托单位:
基于脑微血管内皮细胞模型的毒力岛4在单增李斯特菌CNS炎症中的作用及机制研究
-
批准号:32160834
-
项目类别:地区科学基金项目
-
资助金额:35万元
-
批准年份:2021
-
负责人:马勋
-
依托单位:
胱硫醚-β-合成酶介导小胶质细胞极化致糖皮质激素CNS毒性作用及机制研究
-
批准号:82104317
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:赵瑛
-
依托单位:
生物工程化微泡干扰MAPK通路重编程CNS微环境起始脑胶质瘤免疫检查点抑制剂的应答研究
-
批准号:82102900
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:卢利森
-
依托单位:
百合中“桉树脑盒”挥发物生物合成关键酶基因CNS的功能解析
-
批准号:32002082
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:孔滢
-
依托单位:
新型化合物组合抑制STAT6维持Foxp3-CNS2去甲基化产生稳定的iTreg细胞诱导小鼠肾移植免疫耐受的机制研究
-
批准号:82070773
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2020
-
负责人:陈恕求
-
依托单位:
大气细颗粒物通过NF-κB/LBP-9信号通路诱导小胶质细胞激活加剧CNS脱髓鞘损伤的作用机制研究
-
批准号:82071396
-
项目类别:面上项目
-
资助金额:55.0万元
-
批准年份:2020
-
负责人:张媛
-
依托单位:
环状RNA介导CNS1S1基因影响热应激奶牛乳腺αs1-casein合成及机制研究
-
批准号:32072714
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2020
-
负责人:孙加节
-
依托单位: