A practical tutorial on bagging and boosting based ensembles for machine learning: Algorithms, software tools, performance study, practical perspectives and opportunities

A practical tutorial on bagging and boosting based ensembles for machine learning: Algorithms, software tools, performance study, practical perspectives and opportunities
复制标题

DOI:
10.1016/j.inffus.2020.07.007
复制
发表时间:
2020-12-01
期刊:
影响因子:
18.6
通讯作者:
Herrera, Francisco
Herrera, Francisco
中科院分区:
计算机科学1区
文献类型:
--
作者:
Gonzalez, Sergio;Garcia, Salvador;Herrera, Francisco

文献摘要

被引文献

相似文献

集成,尤其是决策树集成,是机器学习中最流行和最成功的技术之一。最近,基于集成的提案数量稳步增长。因此,有必要确定哪些算法适合特定问题。在本文中,我们的目标是帮助从业者根据他们的问题特征和工作流程选择最佳的集成技术。为此,我们修改了最著名的装袋和增强算法及其软件工具。这些集成在文献中的变体和改进中进行了详细描述。他们的在线可用软件工具会根据实施的版本和功能进行审查。它们根据支持的编程语言和计算范例进行分类。对 14 种不同的基于 bagging 和 boosting 的集成(包括 XGBoost、LightGBM 和随机森林)的性能进行了预测能力和效率方面的实证分析。该比较是在相同的软件环境下使用 76 个不同的分类任务进行的。它们的预测能力通过各种场景进行评估,例如标准多类问题、具有分类特征和大数据的场景。这些方法的效率是通过相当大的数据集进行分析的。集成学习还揭示了一些实用的观点和机会。
Ensembles, especially ensembles of decision trees, are one of the most popular and successful techniques in machine learning. Recently, the number of ensemble-based proposals has grown steadily. Therefore, it is necessary to identify which are the appropriate algorithms for a certain problem. In this paper, we aim to help practitioners to choose the best ensemble technique according to their problem characteristics and their workflow. To do so, we revise the most renowned bagging and boosting algorithms and their software tools. These ensembles are described in detail within their variants and improvements available in the literature. Their online-available software tools are reviewed attending to the implemented versions and features. They are categorized according to their supported programming languages and computing paradigms. The performance of 14 different bagging and boosting based ensembles, including XGBoost, LightGBM and Random Forest, is empirically analyzed in terms of predictive capability and efficiency. This comparison is done under the same software environment with 76 different classification tasks. Their predictive capabilities are evaluated with a wide variety of scenarios, such as standard multi-class problems, scenarios with categorical features and big size data. The efficiency of these methods is analyzed with considerably large data-sets. Several practical perspectives and opportunities are also exposed for ensemble learning.