Tabular data: Deep learning is not all you need

Tabular data: Deep learning is not all you need
复制标题

DOI:
10.1016/j.inffus.2021.11.011
复制
发表时间:
2021-12-10
期刊:
影响因子:
18.6
通讯作者:
Armon, Amitai
Armon, Amitai
中科院分区:
计算机科学1区
文献类型:
--
作者:
Shwartz-Ziv, Ravid;Armon, Amitai

文献摘要

被引文献

相似文献

解决现实数据科学问题的一个关键因素是选择要使用的模型类型。树集成模型(如XGBoost)通常被推荐用于表格数据的分类和回归问题。然而,最近提出了几种针对表格数据的深度学习模型,声称在某些用例中性能优于XGBoost。本文通过在各种数据集上将新的深度模型与XGBoost进行严格比较,探讨这些深度模型是否应该成为表格数据的推荐选项。除了系统地比较它们的性能外,我们还考虑了它们所需的调整和计算。我们的研究表明,XGBoost在数据集上优于这些深度模型,包括提出深度模型的论文中使用的数据集。我们还证明了XGBoost需要更少的调整。从积极的方面来看,我们表明,深度模型和XGBoost的集成在这些数据集上的表现比单独的XGBoost更好。
A key element in solving real-life data science problems is selecting the types of models to use. Tree ensemble models (such as XGBoost) are usually recommended for classification and regression problems with tabular data. However, several deep learning models for tabular data have recently been proposed, claiming to outperform XGBoost for some use cases. This paper explores whether these deep models should be a recommended option for tabular data by rigorously comparing the new deep models to XGBoost on various datasets. In addition to systematically comparing their performance, we consider the tuning and computation they require. Our study shows that XGBoost outperforms these deep models across the datasets, including the datasets used in the papers that proposed the deep models. We also demonstrate that XGBoost requires much less tuning. On the positive side, we show that an ensemble of deep models and XGBoost performs better on these datasets than XGBoost alone.