DISPERSE: A Decentralized Architecture for Content Replication Resilient to Node Failures

DISPERSE: A Decentralized Architecture for Content Replication Resilient to Node Failures
复制标题

DOI:
10.1109/tnsm.2019.2936425
复制
发表时间:
2020-03
影响因子:
5.3
通讯作者:
S. Anand;Ding Ding-Ding;Paolo Gasti;M. O'Neal;M. Conti;K. Balagani
S. Anand;Ding Ding-Ding;Paolo Gasti;M. O'Neal;M. Conti;K. Balagani
中科院分区:
计算机科学2区
文献类型:
--
作者:
S. Anand;Ding Ding-Ding;Paolo Gasti;M. O'Neal;M. Conti;K. Balagani

文献摘要

相似文献

本文介绍了DISPERSE,一个分布式的可扩展的体系结构,提供内容和服务,通过位置无关的存储和复制的内容,对节点故障的恢复能力。当前的内容分发网络(CDN)至少在某种程度上具有集中式结构,因此容易受到单点故障的影响。DISPERSE通过实现完全分散的结构来解决这一限制。DISPERSE是一个两层架构:第一层(前端层)公开服务(例如,Web、SFTP);第二层(后端层)提供内容和应用程序状态的可靠分布式存储。DISPERSE后端层中的内容作为命名数据网络(NDN)内容对象进行存储和交换。这使得DISPERSE能够实现细粒度、独立于位置、完全分散的内容复制机制。我们验证了两个节点故障的情况下,分散的性能。在第一个场景中,内容可以存储在任何DISPERSE节点中,并且所有节点都有同样的可能发生故障。在这种情况下,我们使用非线性优化技术来确定在可用性和延迟约束下的最佳内容副本数量。在第二种情况下,不同的节点以不同的概率发生故障,内容根据其值、节点故障概率和资源可用性存储在节点中。这种情况下解决的最小成本流问题的一个实例。我们的研究结果表明,DISPERSE减少了五个数量级的内容检索失败相比,常见的CDN实现,没有显着增加内容检索延迟。此外,数值结果表明,DISPERSE提高了内容可用性的一个因素的1.3\times- 2.3\times $时,部署的最小成本流算法。
This paper introduces DISPERSE, a distributed scalable architecture for delivery of content and services that provides resilience against node failure through location-independent storage and replication of content. Current content delivery networks (CDNs) have, at least to some degree, a centralized structure thus susceptible to a single point of failure. DISPERSE addresses this limitation by implementing a fully de-centralized structure. DISPERSE is a two-layer architecture: the first layer (front-end layer) exposes services (e.g., Web, SFTP) to clients; the second layer (back-end layer) provides reliable distributed storage of content and application state. Content in DISPERSE’s back-end layer is stored and exchanged as Named Data Network (NDN) content objects. This allows DISPERSE to implement fine-grained, location-independent, fully decentralized content replication mechanisms. We validate the performance of DISPERSE under two node failure scenarios. In the first scenario, content can be stored in any DISPERSE node, and all nodes are equally likely to fail. In this scenario, we use non-linear optimization techniques to determine the optimal number of content copies under availability and latency constraints. In the second scenario, different nodes fail with different probabilities, and content is stored in nodes according to its value, node failure probability, and resource availability. This scenario is addressed as an instance of the minimum cost flow problem. Our results show that DISPERSE reduces the failure of content retrieval by five orders of magnitude compared to common CDN implementations, without significantly increasing content retrieval delay. Further, numerical results show that DISPERSE improves content availability by a factor of $1.3\times - 2.3\times $ when deploying the minimum cost flow algorithm.