PhD Proposal: Efficient Traffic Detection, Scheduling and Transmission in High-Speed Networks
PhD Proposal: Efficient Traffic Detection, Scheduling and Transmission in High-Speed Networks
批准号:
2590767
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
越来越多的公司建立大型数据中心来提供在线服务,包括网络搜索、网络游戏等。为了支持这些业务,在高速网络中采用了多根树拓扑结构,它利用源主机和目的主机之间的多条路径来保证高带宽。一般来说,高速网络中的流量遵循重尾分布,即10%的流量承载大约90%的数据,而大约90%的流量产生10%的总需求。大多数短流是延迟敏感的。在没有事先区分的情况下对它们进行调度将使它们容易经历长时间的排队延迟和数据包重新排序,这将增加它们的总体完成时间并降低应用程序性能。因此,在大规模高速网络中,准确地检测出线路速率下的大流量,并对不同路径上的流量进行适当的调度,以达到低流量竞争时间是至关重要的。此外,需要新的传输协议来提高高速网络中链路资源的利用率。在这个项目中,我将设计一套鲁棒的流量工程解决方案来改善大流量检测精度问题,减少流量竞争次数:(1)高精度流量检测的数据结构:在目前的大流量检测方法中,面对内存约束时,会天真地替换流量持久性计数器,导致检测精度较低。我计划设计一个新的草图方案,通过概率替换的方法更仔细地替换现有的流指标,同时通过记录流键(例如五元组信息)在更新和查询操作方面保持高吞吐量。(2)具有自适应重路由粒度的灵活流量调度:当前高速网络由于流量动态和链路故障等原因,在多种路径下运行,导致拓扑不对称。对于流调度,现有的流级方法将每个流映射到一条路径,而流级方法仅在出现新流时重新路由流。即使在对称拓扑下,这两种方案的交换粒度都很粗,导致网络资源利用率较低。相反,细粒度机制在不对称拓扑中容易出现严重的数据包重排序,从而导致性能下降。为了减少这种情况下的流量完成时间,我将提出一种新的流量调度机制,根据实时网络状态自适应地调整重路由粒度。(3)流量传输的精确子流调整:传统TCP以广域网的可靠传输为目标,依靠丢包作为拥塞信号。普通TCP不适用于数据中心网络,因为它会降低传输性能,特别是短流。这是因为只要没有丢包,数据包就会存储在交换机缓冲区中,这可能导致长队列,从而导致大的流完成时间。MPTCP使用几个子流来避免数据包重新排序并减少流完成时间。然而,目前的MPTCP设计通常不知道子流的数量,导致带宽利用率不足或频繁超时事件。为了避免长时间的排队延迟,为短流提供低延迟传输,同时保证长流的高吞吐量,我计划设计一种传输机制,利用深度强化学习的力量,学习如何在高速网络中根据实时网络状态灵活调整子流的数量。数据集:对于流量检测,我将使用来自CAIDA的匿名IP跟踪。对于流量调度和传输,我将使用公开的Facebook数据集。预期成果:预期成果将是提交给IEEE INFOCOM, IEEE ICNP或ACM CoNEXT的几篇研究论文。开发的代码将在Github上发布。
英文摘要
More and more companies build large data centers to provide online services, including web search, online gaming, etc. To support these services, multi-rooted tree topologies have been employed in high-speed networks, which utilise multiple paths between source and destination hosts to guarantee high bandwidth. In general, traffic in high-speed networks follows a heavy-tailed distribution, i.e. 10% of flows carry approximately 90% of data, while approximately 90% of them generate 10% of the overall demand. Most short flows are delay-sensitive. Scheduling these without prior discrimination will make them prone to experiencing long queuing delays and packet reordering, which increases their overall completion time and deteriorates application performance. Thus, accurately detecting heavy flows at line rate and scheduling traffic on different paths appropriately, to attain low flow competition times is crucial in large-scale high-speed networks. In addition, new transport protocols are needed to increase the utilization of link resources in such high-speed networks. In this project, I will design a set of robust traffic engineering solutions to ameliorate the heavy flow detection accuracy problem and reduce flow competition times: (1) Data structures for highly-accurate flow detection: In current heavy flow detection approaches, flow persistence counters are naively replaced when facing memory constraints, resulting in low detection accuracy. I plan to design a new sketch scheme to replace incumbent flow indicators more carefully through a probabilistic replacement approach, while maintaining high throughput in terms of update and query operations via recording flow keys (e.g. five-tuple information). (2) Flexible flow scheduling with adaptive rerouting granularity: Current high-speed networks run under diverse paths due to traffic dynamics and link failures, which results in topology asymmetries. For flow scheduling, existing flow-level approaches map each flow to one path, and flowlet-level approaches only reroute flowlets when a new flowlet emerges. Both kinds of schemes suffer from low utilization of network resources due to their coarse switching granularity even under symmetric topologies. In contrast, fine-grained mechanisms are prone to severe packet reordering in asymmetric topologies, leading to performance degradation. To decrease flow completion time in such circumstances, I will propose a new flow scheduling mechanism that adjusts the rerouting granularity adaptively according to the real-time network status. (3) Accurate subflow adjustment for flow transmission: Traditional TCP is aimed at reliable transport in wide-area networks, relying on packet loss as the congestion signal. Vanilla TCP is inappropriate for data center networks, since it deteriorates transmission performance, particularly of short flows. This is because packets will be stored in switch buffers as long as there is no packet loss, which may result in long queues and thus large flow completion times. MPTCP uses several subflows to avert packet reordering and reduce flow completion time. However, current MPTCP designs are usually unaware of the number of subflows, leading to bandwidth under-utilization or frequent timeout events. In order to avoid long queuing delays and provide low-latency transmission for short flows while guaranteeing high throughput for long ones, I plan to design a transmission mechanism that harnesses the power of deep reinforcement learning to learn how to adjust the number of subflows flexibly according to real-time network status in high-speed networks. Datasets: For flow detection, I will utilize anonymized IP traces from CAIDA. For flow scheduling and transmission, I will apply the public Facebook dataset. Expected Outputs: The expected outcome will be several research papers to be submitted to IEEE INFOCOM, IEEE ICNP or ACM CoNEXT. The code developed will be published on Github.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金