REBOUND: defending distributed systems against attacks with bounded-time recovery

REBOUND: defending distributed systems against attacks with bounded-time recovery
复制标题

DOI:
10.1145/3447786.3456257
复制
发表时间:
2021-04
期刊:
Proceedings of the Sixteenth European Conference on Computer Systems
影响因子:
--
通讯作者:
N. Gandhi;Edo Roth;Brian Sandler;Andreas Haeberlen;L. T. Phan
N. Gandhi;Edo Roth;Brian Sandler;Andreas Haeberlen;L. T. Phan
中科院分区:
其他
文献类型:
--
作者:
N. Gandhi;Edo Roth;Brian Sandler;Andreas Haeberlen;L. T. Phan

文献摘要

相似文献

本文展示了如何使用有限时间恢复(BTR)来保护分布式系统免受非崩溃故障和攻击。与许多现有的容错技术不同,BTR 并不试图完全掩盖故障的所有症状;相反,它确保系统在有限的时间内返回到正确的行为。这种较弱的保证就足够了,例如,对于许多网络物理系统来说,其中物理属性(例如惯性和热容量)可以防止快速状态变化,从而限制短暂的未定义行为可能造成的损害。我们提出了一种名为 REBOUND 的算法,可以为拜占庭故障模型提供 BTR。 REBOUND 的工作原理是检测故障,然后重新配置系统以排除故障节点。这支持对故障的非常细粒度的响应:例如,系统可以移动或替换现有任务,或者完全删除不太重要的任务以节省资源。即使大多数节点受到损害,REBOUND 也可以采取有用的操作,并且它需要的冗余比完全容错更少。
This paper shows how to use bounded-time recovery (BTR) to defend distributed systems against non-crash faults and attacks. Unlike many existing fault-tolerance techniques, BTR does not attempt to completely mask all symptoms of a fault; instead, it ensures that the system returns to the correct behavior within a bounded amount of time. This weaker guarantee is sufficient, e.g., for many cyber-physical systems, where physical properties - such as inertia and thermal capacity - prevent quick state changes and thus limit the damage that can result from a brief period of undefined behavior. We present an algorithm called REBOUND that can provide BTR for the Byzantine fault model. REBOUND works by detecting faults and then reconfiguring the system to exclude the faulty nodes. This supports very fine-grained responses to faults: for instance, the system can move or replace existing tasks, or drop less critical tasks entirely to conserve resources. REBOUND can take useful actions even when a majority of the nodes is compromised, and it requires less redundancy than full fault-tolerance.