课题基金 / 基金详情

BIGDATA: Collaborative Research: F: RDMA-Based Datacenter Networks for Online Big Data Applications

BIGDATA: Collaborative Research: F: RDMA-Based Datacenter Networks for Online Big Data Applications
BIGDATA:协作研究:F:用于在线大数据应用的基于 RDMA 的数据中心网络
批准号:
1633412
负责人:
Mithuna Thottethodi
金额:
$57.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2021-08-31

项目摘要

项目成果

Mithuna Thottethodi的其他基金

相似基金

相关文献

中文摘要
翻译
该项目解决了在线大数据(OLBD)应用程序在数据中心网络中实现极低延迟的挑战,这些应用程序是数据中心计算的关键工作负载。远程直接内存访问(Remote Direct Memory Access, RDMA)是传统TCP的一种很有前途的替代方案,它将数据中心网络延迟显著降低了大约一个数量级。然而,采用RDMA带来了两个主要挑战,因为RDMA在拥塞情况下存在性能脆弱性,并且RDMA会浪费内存,或者对典型的OLBD流量造成严重的程序员负担。该项目开发了两种新颖的网络技术——Blitz和RIMA,它们使可扩展的数据中心网络能够实现RDMA的低延迟优势,同时避免了RDMA的缺点(性能脆弱性、程序员负担和内存浪费)。Blitz通过解耦边缘拥塞和网络内拥塞来解决性能脆弱性问题。Blitz使用接收方定向拥塞控制(RDCC)处理边缘拥塞,不像以前的方法,发送方必须从往返时间和/或丢弃的数据包间接推断发送速率。RDCC可以实现准确、快速(在一个往返时间内)的收敛,从而实现更低的延迟和更高的吞吐量。Blitz通过使数据包沿着较长但较少拥塞的路径偏转来处理瞬时网络内拥塞。远程间接内存访问(Remote Indirect Memory Access, RIMA)解决了第二个挑战,它支持响应性的、按需的内存分配,而不是rdma在最坏情况下的主动内存分配,这可以最小化内存占用,而无需程序员的努力。Blitz和RIMA共同为OLBD应用程序提供极低的数据中心网络延迟。该项目将广泛涉及博士、硕士和本科生的跨层研究活动。这些项目的成果将在科学会议上通过出版物广泛传播。
英文摘要
This project addresses the challenge of achieving extreme low latency in datacenter networks for Online Big Data (OLBD) applications which are critical workloads in datacenter computing. Remote Direct Memory Access (RDMA), which is a promising alternative to traditional TCP, significantly reduces datacenter network latencies by about an order-of-magnitude. However, RDMA adoption poses two major challenges as RDMA suffers from performance fragility under congestion, and RDMA incurs either wasted memory or significant programmer burden for typical OLBD traffic. This project develops two novel networking technologies -- Blitz and RIMA -- which enable scalable datacenter networks that achieve the low latency benefits of RDMA while avoiding its drawbacks (performance fragility, programmer burden, and wasted memory).Blitz addresses performance fragility by decoupling edge-congestion and in-network congestion. Blitz handles edge-congestion using receiver-directed congestion control (RDCC) unlike prior approaches where senders have to infer sending rates indirectly from round-trip-times and/or dropped packets. RDCC enables accurate and fast (within-one-round-trip-time) convergence, which leads to lower latency and higher throughput. Blitz handles transient in-network congestion by deflecting packets along longer yet less-congested paths.Remote Indirect Memory Access (RIMA) addresses the second challenge by enabling reactive, on-demand memory allocation as opposed to RDMAs proactive memory allocation for the worst case, which minimizes the memory footprint without programmer effort. Together, Blitz and RIMA enable extreme low datacenter network latency for OLBD applications. The project will extensively involve Ph.D, Masters, and undergraduate students in cross-layer research activities. Results from the projects will be broadly disseminated via publication in scientific conferences.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
Slytherin: Dynamic, Network-Assisted Prioritization of Tail Packets in Datacenter Networks
斯莱特林:数据中心网络中尾部数据包的动态网络辅助优先级排序
DOI: 10.1109/icccn.2018.8487331
发表时间: 2018
期刊: International Conference on Computer Communication and Networks (ICCCN
影响因子: --
作者: [Rezaei, Hamed, Malekpourshahraki, Mojtaba, Vamanan, Balajee]
通讯作者: Vamanan, Balajee
Millipede: Die-Stacked Memory Optimizations for Big Data Machine Learning Analytics
Millipede:用于大数据机器学习分析的芯片堆叠内存优化
DOI: 10.1109/ipdps.2018.00026
发表时间: 2018
期刊: IEEE International Parallel and Distributed Processing Symposium (IPDPS
影响因子: --
作者: [Nitin, ., Thottethodi, Mithuna, Vijaykumar, T. N.]
通讯作者: Vijaykumar, T. N.
FastZ: accelerating gapped whole genome alignment on GPUs
FastZ:在 GPU 上加速有缺口的全基因组比对
DOI: 10.1145/3458817.3476202
发表时间: 2021
期刊: Storage and Analysis
影响因子: --
作者: [Gundabolu, Sree Charan, Vijaykumar, T. N., Thottethodi, Mithuna]
通讯作者: Thottethodi, Mithuna
DOI: 10.1145/3374215
发表时间: 2020-05
期刊: ACM Transactions on Architecture and Code Optimization (TACO)
影响因子: --
作者: [Jiachen Xue;Nvidia T. N. Vijaykumar;Mithuna Thottethodi;T. N. Vijaykumar]
通讯作者: Jiachen Xue;Nvidia T. N. Vijaykumar;Mithuna Thottethodi;T. N. Vijaykumar
6
    CSR: Small: SmartEdge for Low Latency and Consistent Mobile Web Applications
    • 批准号:
      1618921
    • 项目类别:
      Standard Grant
    • 资助金额:
      $50.0万
    • 财政年份:
      2016
    • 负责人:
      Mithuna Thottethodi
    • 依托单位:
    CAREER: Cross-Layer Schemes For Flexible Resource Sharing in Multicore Systems
    • 批准号:
      0644183
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $25.0万
    • 财政年份:
      2007
    • 负责人:
      Mithuna Thottethodi
    • 依托单位:
    Performance Models and Systems Optimization for Disk-Bound Applications
    • 批准号:
      0621457
    • 项目类别:
      Standard Grant
    • 资助金额:
      $0.0万
    • 财政年份:
      2006
    • 负责人:
      Mithuna Thottethodi
    • 依托单位:
    CPA: Reconfigurable On-chip Network Design Framework for Fault-Tolerance and Performance
    • 批准号:
      0541385
    • 项目类别:
      Standard Grant
    • 资助金额:
      $25.0万
    • 财政年份:
      2006
    • 负责人:
      Mithuna Thottethodi
    • 依托单位:
    海外基金