Concerto: Leveraging Ensembles for Timely, Accurate Model Training Over Voluminous Datasets
Concerto: Leveraging Ensembles for Timely, Accurate Model Training Over Voluminous Datasets
复制标题
DOI:
10.1109/bdcat50828.2020.00024
复制
发表时间:
2020-12
期刊:
影响因子:
--
通讯作者:
Walid Budgaga;Matthew Malensek;S. Pallickara;S. Pallickara
中科院分区:
文献类型:
--
作者:
Walid Budgaga;Matthew Malensek;S. Pallickara;S. Pallickara
As data volumes increase, there is a pressing need to make sense of the data in a timely fashion. Voluminous datasets are often multidimensional with individual data points representing a vector of features. Data scientists fit models to the data — using all features or a subset thereof — and then use these models to inform their understanding of phenomena or make predictions. The performance of these analytical models is assessed based on their accuracy and ability to generalize on unseen data. Several frameworks exist for drawing insights from voluminous datasets, but have limited scalability (which leads to prolonged training times), poor resource utilization, narrow applicability across problem domains, and insufficient support for combining diverse model fitting algorithms.In this study, we describe our methodology for scalable supervised learning over voluminous datasets. The methodology explores the effect of controlled partitioning of the feature space, as well as how analytical models can be combined to preserve accuracy. Rather than build a single, all-encompassing model, we enable practitioners to construct an ensemble of models that are trained independently in parallel over different portions of the data space. This can provide faster training times and increased prediction accuracy overall; our empirical benchmarks demonstrate the suitability of our approach using real-world data.