Collaborative Research: CNS Core: Small: Understanding Per-Hop Flow Control
Collaborative Research: CNS Core: Small: Understanding Per-Hop Flow Control
批准号:
2006827
负责人:
Mohammad Alizadeh
金额:
$25.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2022-09-30
中文摘要
本研究涉及如何最好地管理数据中心网络资源的争用。数据中心是计算机行业增长最快的领域之一,网络将数据中心中的计算机连接起来,使它们能够进行通信。就像当太多人试图同时使用道路时,道路可能会变得拥堵一样,当太多应用程序试图同时发送数据时,数据中心网络可能会变得拥堵。今天的大多数网络使用端到端控制机制--当网络变得拥塞时,它会向计算机发回信号以降低速度。更快的网络似乎会有所帮助,但事实恰恰相反--网络通信量也在迅速增加,在反馈机制开始控制流量之前,可以发送更多的数据。该项目(华盛顿大学和麻省理工学院研究人员的合作项目)旨在探索一种不同的方法,在网络内、网络交换机之间逐跳进行反馈,只针对那些发送太快的应用程序。数据中心拥塞控制的挑战包括快速增长的工作负载需求、越来越快的链路、较小的平均传输大小、极大的突发流量和有限的交换缓冲区容量。现有的端到端拥塞控制系统在这些设置中远远不是最优的,这对于延迟敏感型应用尤其明显。许多数据中心运营商通过使用优先级和/或以非常低的平均利用率运行网络来进行补偿,但这增加了成本,但并未完全解决问题。本研究试图了解基于逐跳流量控制的数据中心网络拥塞控制替代方法的优势和局限性。这项研究将(I)开发一个理论框架来量化这两种不同方法之间的差异,(Ii)展示在现代可编程数据中心网络交换机上的实际实施,以及(Iii)了解和开发解决方案,以应对在数据中心中使用每跳流量控制的工程挑战。如果成功,研究将有助于以更低的成本在数据中心内和跨数据中心部署新兴的延迟敏感型应用程序,以应对突发性流量模式和行业中正在开发的新兴的超高带宽网络。数据中心网络技术发展迅速,因此本研究的一个关键方面是开发材料,帮助培训本科生和研究生,以应对延迟敏感型应用程序对数据中心网络构成的挑战。项目网站https://www.cs.washington.edu/homes/tom/backpressure/,包含所有项目论文、演示文稿、源代码、模拟、实验结果和教材的副本。随着项目的进展,额外的材料将被放置在那里,并将在项目完成后至少保留十年。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This research concerns how best to manage contention for data center network resources. Data centers are among the fastest growing segment of the computer industry, and networks connect the computers in a data center to allow them to communicate. Just as roads can become congested when too many people try to use them at the same time, data center networks can become congested when too many applications try to send data at the same time. Most networks today use an end to end control mechanism - as the network becomes congested, it sends signals back to the computers to slow down. It might seem that faster networks would help, but the opposite is true - the amount of network communication is also rapidly increasing, and more data can be sent before the feedback mechanism can kick in to control traffic. This project (a collaborative project between investigators at the University of Washington and Massachusetts Institute of Technology) is to explore a different approach, where feedback occurs within the network, hop-by-hop between network switches, and just for those applications that are sending too fast.The challenges for congestion control for data centers include rapidly increasing workload demand, ever faster links, small average transfer sizes, extremely bursty traffic, and limited switch buffer capacity. Existing end-to-end congestion control systems are far from optimal in these settings, and this is particularly noticeable for latency-sensitive applications. Many data center operators compensate by using priorities and/or running their networks at very low average utilization, but this raises costs without fully solving the problem. This research attempts to understand the benefits and limits of an alternative approach to congestion control for data center networks, based on per-hop flow control. The research will (i) develop a theoretical framework to quantify the difference between the two different approaches, (ii) demonstrate a practical implementation on modern programmable data center network switches, and (iii) understand and develop solutions for the engineering challenges of using per-hop flow control in data centers.If successful, the research will help enable an emerging class of latency-sensitive applications to be deployed within and across data centers and at lower cost, for bursty traffic patterns and emerging very high bandwidth networks being developed in industry. Data center network technologies are rapidly evolving, and so a key aspect of this research is to develop materials to help train undergraduate and graduate students for the challenges that latency-sensitive applications pose for data center networks.The project website, https://www.cs.washington.edu/homes/tom/backpressure/, contains copies of all project papers, presentations, source code, simulations, experimental results, and teaching materials. Additional material will be placed there as the project progresses, and will be maintained for a minimum of ten years after the completion of the project.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Scalable Tail Latency Estimation for Data Center Networks
数据中心网络的可扩展尾部延迟估计
DOI:
--
发表时间:
2023
期刊:
20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23
影响因子:
--
作者:
[Zhao, Kevin, Goyal, Prateesh, Alizadeh, Mohammad, Anderson, Thomas E.]
通讯作者:
Anderson, Thomas E.
Collaborative Research: CNS Core: Medium: Learning to Cache and Caching to Learn in High Performance Caching Systems
-
批准号:1955370
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2020
-
负责人:Mohammad Alizadeh
-
依托单位:
CNS Core: Small: Network Architecture and Routing Protocols for Payment Channel Networks
-
批准号:1910676
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2019
-
负责人:Mohammad Alizadeh
-
依托单位:
CAREER: Data-Driven Network Resource Management Systems
-
批准号:1751009
-
项目类别:Continuing Grant
-
资助金额:$62.8万
-
财政年份:2018
-
负责人:Mohammad Alizadeh
-
依托单位:
NeTS: Small: Collaborative Research: A Fast and Flexible Transport Architecture for High Speed Networks
-
批准号:1617702
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2016
-
负责人:Mohammad Alizadeh
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: