A Comparative Study of Penalized Regression and Machine Learning Algorithms in High Dimensional Scenarios

A Comparative Study of Penalized Regression and Machine Learning Algorithms in High Dimensional Scenarios
复制标题

DOI:
10.1137/22s1538302
复制
发表时间:
2023
期刊:
SIAM Undergraduate Research Online
影响因子:
--
通讯作者:
Connor Shrader;Gabriel Ackall
Connor Shrader;Gabriel Ackall
中科院分区:
其他
文献类型:
--
作者:
Connor Shrader;Gabriel Ackall

文献摘要

相似文献

.随着近年来大数据的流行,对高维数据进行建模和选择重要特征的重要性大大增加。高维数据在基因组解码、罕见疾病识别和环境建模等许多领域都很常见。然而,大多数传统的回归机器学习模型并不是为了处理高维数据或进行变量选择而设计的。在本文中,我们研究了惩罚回归方法的使用,如岭,最小绝对收缩和选择操作,弹性网络,平滑裁剪绝对偏差,和minimax凹罚相比,传统的机器学习模型,如随机森林,XGBoost,和支持向量机。我们使用析因设计方法在540个环境中进行Monte Carlo模拟比较这些模型,因子为响应变量、预测因子数、样本数、信噪比、协方差矩阵和相关强度。我们还比较了不同的模型,使用经验数据来评估其在现实世界中的可行性。我们使用训练和测试均方误差,变量选择准确性,β敏感性和β特异性来评估模型。我们发现,在大多数高维情况下,惩罚回归模型的性能与传统机器学习算法相当。该分析有助于更好地了解每种模型类型的优势和劣势,并为其他研究人员提供参考,根据一系列因素和数据环境,他们应该使用哪些机器学习技术。我们的研究表明,惩罚回归技术应该包括在预测建模的工具箱。
. With the prevalence of big data in recent years, the importance of modeling high dimensional data and selecting important features has increased greatly. High dimensional data is common in many fields such as genome decoding, rare disease identification, and environmental modeling. However, most traditional regression machine learning models are not designed to handle high dimensional data or conduct variable selection. In this paper, we investigate the use of penalized regression meth-ods such as ridge, least absolute shrinkage and selection operation, elastic net, smoothly clipped absolute deviation, and minimax concave penalty compared to traditional machine learning models such as random forest, XGBoost, and support vector machines. We compare these models using factorial design methods for Monte Carlo simulations in 540 environments, with factors being the response variable, number of predictors, number of samples, signal to noise ratio, covariance matrix, and correlation strength. We also compare different models using empirical data to evaluate their viability in real-world scenarios. We evaluate the models using the training and test mean squared error, variable selection accuracy, β -sensitivity, and β -specificity. We found that the performance of penalized regression models is comparable with traditional machine learning algorithms in most high-dimensional situations. The analysis helps to create a greater understanding of the strengths and weaknesses of each model type and provide a reference for other researchers on which machine learning techniques they should use, depending on a range of factors and data environments. Our study shows that penalized regression techniques should be included in predictive modelers’ toolbox.