Concerto: Leveraging Ensembles for Timely, Accurate Model Training Over Voluminous Datasets

Concerto: Leveraging Ensembles for Timely, Accurate Model Training Over Voluminous Datasets
复制标题

DOI:
10.1109/bdcat50828.2020.00024
复制
发表时间:
2020-12
期刊:
2020 IEEE/ACM International Conference on Big Data Computing, Applications and Technologies (BDCAT)
影响因子:
--
通讯作者:
Walid Budgaga;Matthew Malensek;S. Pallickara;S. Pallickara
Walid Budgaga;Matthew Malensek;S. Pallickara;S. Pallickara
中科院分区:
其他
文献类型:
--
作者:
Walid Budgaga;Matthew Malensek;S. Pallickara;S. Pallickara

文献摘要

相似文献

随着数据量的增加,迫切需要及时地理解数据。大量的数据集通常是多维的,单个数据点表示特征向量。数据科学家使用所有特征或其子集将模型拟合到数据中,然后使用这些模型来告知他们对现象的理解或做出预测。这些分析模型的性能是根据其准确性和对未知数据进行概括的能力进行评估的。有几个框架存在从海量数据集的见解,但有限的可扩展性(这导致延长训练时间),资源利用率低,适用性狭窄的问题域,并结合不同的模型拟合algorithm.In这项研究中,我们描述了我们的方法,可扩展的监督学习在海量数据集。该方法探讨了特征空间的控制分区的效果,以及如何结合分析模型以保持准确性。而不是建立一个单一的,包罗万象的模型,我们使从业者能够构建一个模型的集合,这些模型在数据空间的不同部分并行独立训练。这可以提供更快的训练时间和更高的整体预测准确性;我们的经验基准证明了我们使用真实数据的方法的适用性。
As data volumes increase, there is a pressing need to make sense of the data in a timely fashion. Voluminous datasets are often multidimensional with individual data points representing a vector of features. Data scientists fit models to the data — using all features or a subset thereof — and then use these models to inform their understanding of phenomena or make predictions. The performance of these analytical models is assessed based on their accuracy and ability to generalize on unseen data. Several frameworks exist for drawing insights from voluminous datasets, but have limited scalability (which leads to prolonged training times), poor resource utilization, narrow applicability across problem domains, and insufficient support for combining diverse model fitting algorithms.In this study, we describe our methodology for scalable supervised learning over voluminous datasets. The methodology explores the effect of controlled partitioning of the feature space, as well as how analytical models can be combined to preserve accuracy. Rather than build a single, all-encompassing model, we enable practitioners to construct an ensemble of models that are trained independently in parallel over different portions of the data space. This can provide faster training times and increased prediction accuracy overall; our empirical benchmarks demonstrate the suitability of our approach using real-world data.