Optimizing Machine Learning Inference Queries with Correlative Proxy Models

Optimizing Machine Learning Inference Queries with Correlative Proxy Models
复制标题

DOI:
10.14778/3547305.3547310
复制
发表时间:
2022-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhihui Yang;Zuozhi Wang;Yicong Huang;Yao Lu;Chen Li-;X. S. Wang
Zhihui Yang;Zuozhi Wang;Yicong Huang;Yao Lu;Chen Li-;X. S. Wang
中科院分区:
其他
文献类型:
--
作者:
Zhihui Yang;Zuozhi Wang;Yicong Huang;Yao Lu;Chen Li-;X. S. Wang

文献摘要

相似文献

我们考虑加速非结构化数据集上的机器学习 (ML) 推理查询。昂贵的运算符(例如特征提取器和分类器)被部署为用户定义函数(UDF),而经典查询优化技术(例如谓词下推)无法穿透这些运算符。最近的优化方案(例如,概率谓词或 PP)假设查询谓词之间是独立的,为每个谓词离线构建代理模型,并通过在昂贵的 ML UDF 前面注入这些廉价的代理模型来重写新的查询。通过这种方式,不满足查询谓词的不太可能的输入会被提前过滤以绕过 ML UDF。我们表明,在这种情况下强制执行独立性假设可能会导致计划次优。在本文中,我们提出了 CORE,一种查询优化器,可以更好地利用谓词相关性并加速 ML 推理查询。我们的解决方案为新查询在线构建代理模型,并利用分支定界搜索过程来降低构建成本。三个真实文本、图像和视频数据集的结果表明,与 PP 相比,CORE 将查询吞吐量提高了 63%,与按原样运行查询相比,提高了 80%。
We consider accelerating machine learning (ML) inference queries on unstructured datasets. Expensive operators such as feature extractors and classifiers are deployed as user-defined functions (UDFs), which are not penetrable with classic query optimization techniques such as predicate push-down. Recent optimization schemes (e.g., Probabilistic Predicates or PP) assume independence among the query predicates, build a proxy model for each predicate offline, and rewrite a new query by injecting these cheap proxy models in the front of the expensive ML UDFs. In such a manner, unlikely inputs that do not satisfy query predicates are filtered early to bypass the ML UDFs. We show that enforcing the independence assumption in this context may result in sub-optimal plans. In this paper, we propose CORE, a query optimizer that better exploits the predicate correlations and accelerates ML inference queries. Our solution builds the proxy models online for a new query and leverages a branch-and-bound search process to reduce the building costs. Results on three real-world text, image and video datasets show that CORE improves the query throughput by up to 63% compared to PP and up to 80% compared to running the queries as it is.