Joint Progressive Network and Datacenter Recovery After Large-Scale Disasters

Joint Progressive Network and Datacenter Recovery After Large-Scale Disasters
复制标题

DOI:
10.1109/tnsm.2020.2983822
复制
发表时间:
2020-09
影响因子:
5.3
通讯作者:
Sifat Ferdousi;M. Tornatore;F. Dikbiyik;C. Martel;Sugang Xu;Y. Hirota;Y. Awaji;B. Mukherjee
Sifat Ferdousi;M. Tornatore;F. Dikbiyik;C. Martel;Sugang Xu;Y. Hirota;Y. Awaji;B. Mukherjee
中科院分区:
计算机科学2区
文献类型:
--
作者:
Sifat Ferdousi;M. Tornatore;F. Dikbiyik;C. Martel;Sugang Xu;Y. Hirota;Y. Awaji;B. Mukherjee

文献摘要

被引文献

相似文献

影响网络和数据中心(DC)基础设施的大规模灾难可能导致基于云的服务严重中断。在灾后恢复过程中,由于可用的维修资源有限,通常以渐进的方式分阶段进行维修。网元和数据中心的修复顺序会严重影响用户访问重要内容/服务的可达性。我们研究了联合渐进式网络和数据中心恢复,其中网络恢复和数据中心恢复以协调的方式进行,以便用户在每个修复阶段都能访问尽可能多的内容/服务。首先以网络中累积加权内容可达性最大化为目标,解决联合渐进恢复的优化问题,找到网元和DC修复的最优顺序。然后,我们提出了一种可扩展的启发式算法,用于调度网络节点/链路和数据中心的顺序修复。我们的模型假设,在每个修复阶段,有一个相邻链路的网络节点和一个DC可以完全修复;但是,由于有限的资源可用性,可能无法保证完全恢复。因此,我们还提出了一种“资源感知”方法(包括两种资源分配策略,即“选择性分配”和“自适应分配”),该方法根据每个阶段的可用资源考虑元素的完全恢复和部分恢复。我们表明,与网络恢复和数据中心恢复计划独立的不连接渐进恢复方法相比,我们的联合渐进恢复方法在网络中提供了更高的每级内容可达性。
Large-scale disasters affecting both network and datacenter (DC) infrastructures can cause severe disruptions in cloud-based services. During post-disaster recovery, repairs are usually carried out in stages in a progressive manner due to limited repair resource availability. The order in which network elements and DCs are repaired can significantly impact users’ reachability to important contents/services. We investigate joint progressive network and DC recovery in which network recovery and DC recovery are conducted in a coordinated manner such that users have access to the maximum possible amount of contents/services at each repair stage. We first solve the optimization problem of joint progressive recovery to find the optimal sequence of network element and DC repairs with the objective to maximize cumulative weighted content reachability in the network. We then propose a scalable heuristic for scheduling the sequential repair of network nodes/links and DCs. Our model assumes that, at each repair stage, one network node with adjacent links and one DC can be fully repaired; however, full recovery may not be guaranteed due to limited resource availability. Hence, we also propose a “resource-aware” approach (with two resource-allocation strategies, namely “selective allocation” and “adaptive allocation”), which considers both full and partial recovery of elements based on available resources at each stage. We show that, compared to disjoint progressive recovery approach, in which network recovery and DC recovery plans are independent, our joint progressive recovery approach provides significantly higher per-stage content reachability in the network.