REINFORCE: Achieving Efficient Failure Resiliency for Network Function Virtualization-Based Services

REINFORCE: Achieving Efficient Failure Resiliency for Network Function Virtualization-Based Services
复制标题

DOI:
10.1145/3281411.3281441
复制
发表时间:
2018-12
期刊:
IEEE/ACM Transactions on Networking
影响因子:
--
通讯作者:
Sameer G. Kulkarni;Guyue Liu;K. Ramakrishnan;M. Arumaithurai;Timothy Wood;Xiaoming Fu
Sameer G. Kulkarni;Guyue Liu;K. Ramakrishnan;M. Arumaithurai;Timothy Wood;Xiaoming Fu
中科院分区:
其他
文献类型:
--
作者:
Sameer G. Kulkarni;Guyue Liu;K. Ramakrishnan;M. Arumaithurai;Timothy Wood;Xiaoming Fu

文献摘要

被引文献

相似文献

确保基于软件的网络的高可用性(HA)是一项关键的设计功能,有助于在生产网络中采用基于软件的网络功能(NF)。对于NF来说,避免中断和维护关键任务操作非常重要。但是,对关键数据路径上NF的HA支持可能导致不可接受的性能下降。我们提出了REINFORCE,一个集成的框架,以支持NF服务链的有效弹性。REINFORCE包括及时的故障检测和一致的故障转移机制。REINFORCE将状态复制到备用NF(本地和远程),同时强制执行正确性。它通过利用外部同步的概念来最大限度地减少状态传输的数量,并利用机会主义的同步和多缓冲来优化性能。实验结果表明,即使在10 Gbps的线速数据包处理速率下,REINFORCE也能在10 ms内实现跨局域网服务器的链级故障转移,性能开销小于10%,平均延迟仅增加$\sim 400~\mu \text{s}$,最坏情况下延迟小于1 ms。REINFORCE还可以在不到100美元的时间内从同一节点内的软件故障中恢复,在正常操作期间,性能开销不到1%,延迟不到5美元。
Ensuring high availability (HA) for software-based networks is a critical design feature that will help the adoption of software-based network functions (NFs) in production networks. It is important for NFs to avoid outages and maintain mission-critical operations. However, HA support for NFs on the critical data path can result in unacceptable performance degradation. We present REINFORCE, an integrated framework to support efficient resiliency for NF service chains. REINFORCE includes timely failure detection and consistent failover mechanisms. REINFORCE replicates state to standby NFs (local and remote) while enforcing correctness. It minimizes the number of state transfers by exploiting the concept of external synchrony, and leverages opportunistic batching and multi-buffering to optimize performance. Experimental results show that, even at line-rate packet processing (10 Gbps), REINFORCE achieves chain-level failover across servers in a LAN within 10ms, incurring less than 10% performance overhead, and adds average latency only $\sim 400~\mu \text{s}$ , with a worst-case latency of less than 1ms. REINFORCE also recovers from software failures within the same node in less than $100~\mu \text{s}$ , incurring less than 1% performance overhead and adds less than $5~\mu \text{s}$ latency during normal operation.