Exploiting Correlations for Expensive Predicate Evaluation

Exploiting Correlations for Expensive Predicate Evaluation
复制标题

利用相关性进行昂贵的谓词评估

DOI:
10.1145/2723372.2723715
复制
发表时间:
2014
期刊:
Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data
影响因子:
--
通讯作者:
Christopher Ré
Christopher Ré
中科院分区:
--
文献类型:
--
作者:
Manas R. Joglekar;H. Garcia;Aditya G. Parameswaran;Christopher Ré

文献摘要

被引文献

相似文献

用户定义函数(UDF)越来越多地被用来增加查询语言的额外应用程序相关功能。涉及UDF谓词的选择查询往往是昂贵的,无论是在货币成本还是延迟方面。在本文中,我们研究如何有效地评估选择查询与UDF谓词。我们提供了一个家庭的技术,以低成本处理查询,同时满足用户指定的精度和召回的限制。我们的技术适用于各种情况下,包括当元组的选择概率是事先可用的,当此信息是可用的,但嘈杂的,或者当没有这样的先验信息是可用的。我们还将我们的技术推广到更复杂的查询。最后,我们在真实的数据集上测试了我们的技术,并表明它们在UDF评估中实现了高达80\%$的显着节省,同时只会导致精度的小幅降低。
User Defined Function(UDFs) are used increasingly to augment query languages with extra, application dependent functionality. Selection queries involving UDF predicates tend to be expensive, either in terms of monetary cost or latency. In this paper, we study ways to efficiently evaluate selection queries with UDF predicates. We provide a family of techniques for processing queries at low cost while satisfying user-specified precision and recall constraints. Our techniques are applicable to a variety of scenarios including when selection probabilities of tuples are available beforehand, when this information is available but noisy, or when no such prior information is available. We also generalize our techniques to more complex queries. Finally, we test our techniques on real datasets, and show that they achieve significant savings in UDF evaluations of up to $80\%$, while incurring only a small reduction in accuracy.