TerseCades: Efficient Data Compression in Stream Processing

TerseCades: Efficient Data Compression in Stream Processing
复制标题

DOI:
--
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
Gennady Pekhimenko;Chuanxiong Guo;Myeongjae Jeon;Peng Huang;Lidong Zhou
Gennady Pekhimenko;Chuanxiong Guo;Myeongjae Jeon;Peng Huang;Lidong Zhou
中科院分区:
其他
文献类型:
--
作者:
Gennady Pekhimenko;Chuanxiong Guo;Myeongjae Jeon;Peng Huang;Lidong Zhou

文献摘要

被引文献

相似文献

这项工作是第一个系统的调查流处理与数据压缩:我们不仅确定了一组因素,影响压缩的好处和开销,但也证明了压缩可以有效的流处理,无论是在处理能力,在更大的窗口和吞吐量。这是通过一系列(i)对流引擎本身进行优化以消除低效率的主要来源,这导致吞吐量的数量级改进(ii)优化以降低压缩(解压缩)成本,包括硬件加速,以及(iii)允许直接执行压缩数据的新技术,这导致吞吐量进一步提高50%。我们的评估是在云分析和故障排除的几个真实场景中进行的,包括微基准测试和生产流处理系统。
This work is the first systematic investigation of stream processing with data compression: we have not only identified a set of factors that influence the benefits and overheads of compression, but have also demonstrated that compression can be effective for stream processing, both in the ability to process in larger windows and in throughput. This is done through a series of (i) optimizations on a stream engine itself to remove major sources of inefficiency, which leads to an order-of-magnitude improvement in throughput (ii) optimizations to reduce the cost of (de)compression, including hardware acceleration, and (iii) a new technique that allows direct execution on compressed data, that leads to a further 50% improvement in throughout. Our evaluation is performed on several real-world scenarios in cloud analytics and troubleshooting, with both microbenchmarks and production stream processing systems.