Transparent caching with strong consistency in dynamic content web sites

Transparent caching with strong consistency in dynamic content web sites
复制标题

动态内容网站中具有强一致性的透明缓存

DOI:
--
复制
发表时间:
2005
期刊:
International Conference on Supercomputing
影响因子:
--
通讯作者:
E. Cecchet
E. Cecchet
中科院分区:
--
文献类型:
--
作者:
C. Amza;Gokul Soundararajan;E. Cecchet

文献摘要

被引文献

相似文献

我们考虑一个集群体系结构,其中动态内容是由一个数据库后端和Web和应用程序服务器前端的集合。我们研究了透明查询缓存对这样一个集群的性能的影响。透明性要求缓存条目在写入时失效。我们从一个粗粒度的表级自动失效缓存开始。基于观察到的工作负载特征,我们在列的更细粒度上通过必要的依赖跟踪和无效化来增强该高速缓存。最后,我们通过查询结果的完全覆盖和部分覆盖来减少无效的未命中惩罚。在系统设计方面,查询缓存可以位于数据库后端,专用机器上,前端,或其组合。本文评估了不同的缓存设计和该高速缓存的位置使用TPC-W benchmark.Our实验表明,我们的透明查询缓存提高性能非常大的吞吐量和响应时间的整体相比,基线基于表的无效方案的1.5倍和4.2倍。这一最终结果的一个重要贡献者,我们的优化,以减少未命中的惩罚,通过从该高速缓存查询结果的完全和部分覆盖检测提高响应时间高达2.9倍相比,细粒度的基于列的无效缓存单独。因此,在我们的优化中,更高命中率的好处超过了额外处理的成本。结果在定位该高速缓存的位置方面不太明确。改变该高速缓存位置和缓存数量时的性能差异很小。
We consider a cluster architecture in which dynamic content is generated by a database back-end and a collection of Web and application server front-ends. We study the effect of transparent query caching on the performance of such a cluster. Transparency requires that cached entries be invalidated as a result of writes. We start with a coarse-grain table-level automatic invalidation cache. Based on observed workload characteristics, we enhance the cache with the necessary dependency tracking and invalidations at the finer granularity of columns. Finally we reduce the miss penalty of invalidations through full and partial coverage of query results.In terms of system design, a query cache may be located at the database back-end, on dedicated machines, on the front-ends, or on a combination thereof. This paper evaluates the tradeoffs of the different cache designs and the cache location using the TPC-W benchmark.Our experiments show that our transparent query cache improves performance very substantially by up to a factor of 1.5 in throughput and 4.2 in response time overall compared to the baseline table-based invalidation scheme. An important contributor to this end result, our optimization for reducing the miss penalty through full and partial coverage detection of query results from the cache improves response time by up to a factor of 2.9 compared to a cache with fine-grained column-based invalidations alone. Thus, the benefits of the higher hit ratio in our optimizations outweigh the costs of additional processing. The results are less clear-cut in terms of where to locate the cache. Performance differences when varying the cache location and the number of caches are small.