Cutting Latency Tail: Analyzing and Validating Replication without Canceling

Cutting Latency Tail: Analyzing and Validating Replication without Canceling
复制标题

DOI:
10.1109/tpds.2017.2706268
复制
发表时间:
2017-11
影响因子:
5.3
通讯作者:
Z. Qiu;Juan F. Pérez;R. Birke;L. Chen;P. Harrison
Z. Qiu;Juan F. Pérez;R. Birke;L. Chen;P. Harrison
中科院分区:
计算机科学2区
文献类型:
--
作者:
Z. Qiu;Juan F. Pérez;R. Birke;L. Chen;P. Harrison

文献摘要

被引文献

相似文献

软件应用程序中的响应时间可变性会严重降低用户体验的质量。为了减少这种可变性,请求复制成为一种有效的解决方案,它为每个请求生成多个副本,并使用第一个副本的结果来完成。大多数以前的研究主要集中在实现副本消除的系统的平均延迟上,即,一旦第一个请求完成,则取消请求的所有副本。相反,我们开发的模型,以获得副本取消可能过于昂贵或不可行的实施,如在“快速”系统,如Web服务,或在遗留系统的系统的响应时间分布。此外,我们引入了一种新的服务模型,明确考虑相关性的请求副本的处理时间,并设计了一个有效的算法来参数化模型从真实的数据。在MATLAB基准测试和三层Web应用程序(MediaWiki)上进行的广泛评估显示出显著的准确性,例如,7(4%)基准测试(分别为MediaWiki)第99百分位响应时间的平均错误,其请求以秒(分别为毫秒)的数量级执行。因此,在各种各样的系统场景下,通过这种精确的定量分析,可以深入了解最佳复制级别。
Response time variability in software applications can severely degrade the quality of the user experience. To reduce this variability, request replication emerges as an effective solution by spawning multiple copies of each request and using the result of the first one to complete. Most previous studies have mainly focused on the mean latency for systems implementing replica cancellation, i.e., all replicas of a request are canceled once the first one finishes. Instead, we develop models to obtain the response-time distribution for systems where replica cancellation may be too expensive or infeasible to implement, as in “fast” systems, such as web services, or in legacy systems. Furthermore, we introduce a novel service model to explicitly consider correlation in the processing times of the request replicas, and design an efficient algorithm to parameterize the model from real data. Extensive evaluations on a MATLAB benchmark and a three-tier web application (MediaWiki) show remarkable accuracy, e.g., 7 (4 percent) average error on the 99th percentile response time for the benchmark (respectively, MediaWiki), the requests of which execute in the order of seconds (respectively, milliseconds). Insights into optimal replication levels are thereby gained from this precise quantitative analysis, under a wide variety of system scenarios.