RobinHood: Tail Latency Aware Caching - Dynamic Reallocation from Cache-Rich to Cache-Poor

RobinHood: Tail Latency Aware Caching - Dynamic Reallocation from Cache-Rich to Cache-Poor
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
--
影响因子:
--
通讯作者:
Daniel S. Berger;Benjamin Berg;T. Zhu;S. Sen;Mor Harchol-Balter
Daniel S. Berger;Benjamin Berg;T. Zhu;S. Sen;Mor Harchol-Balter
中科院分区:
其他
文献类型:
--
作者:
Daniel S. Berger;Benjamin Berg;T. Zhu;S. Sen;Mor Harchol-Balter

文献摘要

相似文献

尾部潜伏期在面向用户的Web服务中非常重要。但是,保持低尾部潜伏期是具有挑战性的,因为对Web应用程序服务器的单个请求会引起对复杂,多样化的后端服务(数据库,推荐系统,AD系统等)的多个查询。在所有查询都完成之前,请求还不完整。我们分析了一个Microsoft生产系统,并发现后端查询潜伏期在整个后端和随着时间的流逝之间的变化两个以上的数量级,从而产生了高要求的尾巴潜伏期。我们提出了一种新颖的解决方案,以保持低要求尾巴潜伏期:重新利用现有的缓存,以减轻后端延迟可变性的影响,而不仅仅是缓解流行数据。我们的解决方案,Robinhood,动态地重新关注了从高速缓存的高速缓存资源(不影响请求尾部潜伏期的后端)到缓存贫乏(影响请求尾部潜伏期的后端)。我们在具有20个不同后端系统的50服务器群集上使用生产轨迹评估了Robinhood。令人惊讶的是,我们发现,即使工作组比缓存大小要大得多,Robinhood也可以直接解决尾部潜伏期。在有负载峰值的情况下,罗比尼(Robinhood)在99.7%的时间内达到了150毫秒的P99进球,而下一个最佳政策仅在70%的时间内达到了这一目标。
Tail latency is of great importance in user-facing web services. However, maintaining low tail latency is challenging, because a single request to a web application server results in multiple queries to complex, diverse backend services (databases, recommender systems, ad systems, etc.). A request is not complete until all of its queries have completed. We analyze a Microsoft production system and find that backend query latencies vary by more than two orders of magnitude across backends and over time, resulting in high request tail latencies. We propose a novel solution for maintaining low request tail latency: repurpose existing caches to mitigate the effects of backend latency variability, rather than just caching popular data. Our solution, RobinHood, dynamically reallocates cache resources from the cache-rich (backends which don't affect request tail latency) to the cache-poor (backends which affect request tail latency). We evaluate RobinHood with production traces on a 50- server cluster with 20 different backend systems. Surprisingly, we find that RobinHood can directly address tail latency even if working sets are much larger than the cache size. In the presence of load spikes, RobinHood meets a 150ms P99 goal 99.7% of the time, whereas the next best policy meets this goal only 70% of the time.