Trading off Accuracy for Speed in PowerDrill

Trading off Accuracy for Speed in PowerDrill
复制标题

在 PowerDrill 中以精度换取速度

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
Thomas Hofmann
Thomas Hofmann
中科院分区:
--
文献类型:
--
作者:
A. Hall;A. Tudorica;F. Buruiana;R. Hofmann;Silviu;Thomas Hofmann

文献摘要

被引文献

相似文献

内存中的列存储使交互式分析在许多大数据场景中变得可行。在本文中,我们研究了两种正交的方法来优化性能的代价是一个可接受的损失的准确性。这两种方法都可以作为现有数据库引擎的外部包装器来实现,因此它们应该很容易应用于其他系统。对于第一个优化,我们表明,内存是执行查询速度的限制因素,因此探索提高内存效率的可能性。与标准压缩算法相比,我们采用了数据草图背后的一些理论,将最大表中特别昂贵的字段的大小减少了4.5倍。这在PowerDrill中节省了37%的总内存,并为具有昂贵字段的查询结果引入了0.4%的第90百分位相对误差。我们还评估了使用采样的准确性的影响,并提出了一个简单的启发式注释个别结果值准确(或不准确)。在我们的真实的生产系统中的用户行为的测量的基础上,我们表明,这些估计是必不可少的解释中间结果之前,最终的结果。对于大量查询,这有效地将第95个延迟百分比从30秒降低到4秒。
In-memory column-stores make interactive analysis feasible for many big data scenarios. In this paper we investigate two orthogonal approaches to optimize performance at the expense of an acceptable loss of accuracy. Both approaches can be implemented as outer wrappers around existing database engines and so they should be easily applicable to other systems. For the first optimization we show that memory is the limiting factor in executing queries at speed and therefore explore possibilities to improve memory efficiency. We adapt some of the theory behind data sketches to reduce the size of particularly expensive fields in our largest tables by a factor of 4.5 when compared to a standard compression algorithm. This saves 37% of the overall memory in PowerDrill and introduces a 0.4% relative error in the 90th percentile for results of queries with the expensive fields. We additionally evaluate the effects of using sampling on accuracy and propose a simple heuristic for annotating individual result-values as accurate (or not). Based on measurements of user behavior in our real production system, we show that these estimates are essential for interpreting intermediate results before final results are available. For a large set of queries this effectively brings down the 95th latency percentile from 30 to 4 seconds.