Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse

Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse
复制标题

DOI:
10.14778/3236187.3236199
复制
发表时间:
2018-07
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Sona Hasani;Saravanan Thirumuruganathan;Abolfazl Asudeh;Nick Koudas;Gautam Das
Sona Hasani;Saravanan Thirumuruganathan;Abolfazl Asudeh;Nick Koudas;Gautam Das
中科院分区:
其他
文献类型:
--
作者:
Sona Hasani;Saravanan Thirumuruganathan;Abolfazl Asudeh;Nick Koudas;Gautam Das

文献摘要

相似文献

机器学习已经成为复杂分析处理的重要工具。数据通常存储在具有多个维度层次结构的大型数据仓库中。通常,用于构建ML模型的数据在OLAP层次结构上对齐,例如位置或时间。在本文中,我们研究了利用模型物化和重用的概念,从以前构建的ML模型中有效地为新查询构建近似ML模型的可行性。例如,如果已经有了每个季度的ML模型,那么是否可以为2017年的数据构建一个近似的ML模型?我们提出了可以支持各种ML模型的算法,例如用于分类的广义线性模型,沿着用于聚类的K-Means和高斯混合模型。我们提出了一个基于成本的优化框架,该框架可以识别适当的ML模型,以便在查询时进行联合收割机组合,并在真实世界和合成数据集上进行广泛的实验。我们的研究结果表明,我们的框架可以支持ML模型上的分析查询,具有上级性能,在非常大的数据集上实现了几个数量级的显着加速。
Machine learning has become an essential toolkit for complex analytic processing. Data is typically stored in large data warehouses with multiple dimension hierarchies. Often, data used for building an ML model are aligned on OLAP hierarchies such as location or time. In this paper, we investigate the feasibility of efficiently constructing approximate ML models for new queries from previously constructed ML models by leveraging the concepts of model materialization and reuse . For example, is it possible to construct an approximate ML model for data from the year 2017 if one already has ML models for each of its quarters? We propose algorithms that can support a wide variety of ML models such as generalized linear models for classification along with K-Means and Gaussian Mixture models for clustering. We propose a cost based optimization framework that identifies appropriate ML models to combine at query time and conduct extensive experiments on real-world and synthetic datasets. Our results indicate that our framework can support analytic queries on ML models, with superior performance, achieving dramatic speedups of several orders in magnitude on very large datasets.