Group Lasso with Checkpoints Selection for Biological Data Regression

Group Lasso with Checkpoints Selection for Biological Data Regression
复制标题

DOI:
10.1109/smc53992.2023.10394634
复制
发表时间:
2023-10
期刊:
2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC)
影响因子:
--
通讯作者:
Huixin Zhan;Yifan Wang
Huixin Zhan;Yifan Wang
中科院分区:
其他
文献类型:
--
作者:
Huixin Zhan;Yifan Wang

文献摘要

相似文献

生物学数据的一些独特特征是(1)它们始终是高维度和低样本尺寸(HDLS),并且(2)数据分布发生变化,例如类,分布和协变量的不平衡等等。在本文中,我们建议使用检查点选项(GL_CSE)Algorithm的组套件,以解决这两个问题。为了解决第一个问题,我们利用针对HDLSS数据量身定制的组LASSO回归模型在预定义的特征组上执行功能选择,从而减轻过度拟合并在小组正交重新分析下不变。为了解决第二个问题,我们提出了检查点选择方法,以提取重要的模型检查点,同时通过两个提议的指标对小组套索进行培训,即训练和验证功能之间的平均KL差异以及训练和验证功能之间协方差矩阵的Frobenius错误。这两个指标旨在选择模型检查点,而训练和验证功能之间的漂移最少。我们的实验结果表明,与其他基线方法相比,我们提出的GL_CSE算法在MSE和R2测量方面取得了更好的性能。具体而言,在生物年龄数据集上,我们的GL_CSE方法分别为MSE和R2测量值分别达到0.8799和0.9883。此外,我们还表明,我们提出的检查点选择方法的性能优于常规K折交叉验证。具体而言,在生物年龄数据集上,GL_CSE(Q2)分别达到0.9045 MSE和0.9880 R2,这表现优于常规K折叠结果,即分别为1.0612 MSE和0.9871 R2。
Some unique characteristics of biological data are (1) that they are always High-Dimension and Low-Sample-Size (HDLSS) and (2) there are changes in the data distribution, such as an imbalance in classes, distribution and covariate shifts, etc. In this paper, we propose a Group Lasso with Checkpoints SElection (GL_CSE) algorithm to tackle both issues. To address the first issue, we utilize a group Lasso regression model tailored for HDLSS data to perform feature selection on predefined groups of features, alleviating overfitting and being invariant under group-wise orthogonal reparameterizations. To address the second issue, we propose the checkpoint selection method to extract important model checkpoints while training on group Lasso via two proposed metrics, i.e., the average KL-divergence between training and validation features and the Frobenius error of the covariance matrices between training and validation features. Both metrics aim to select model checkpoints with minimal drifts between the training and validation features. The results of our experiments indicate that our proposed GL_CSE algorithm achieves better performance compared to other baseline methods in terms of the MSE and R2measurements. Specifically, on the biological age dataset, our GL_CSE method achieves 0.8799 and 0.9883 for the MSE and R2 measurements, respectively. Additionally, we also show that our proposed checkpoint selection method performs better than regular K-fold cross-validation. Specifically, on the biological age dataset, GL_CSE (Q2) achieves 0.9045 MSE and 0.9880 R2, respectively, which outperforms the regular K-fold cross-validation results, i.e., 1.0612 MSE and 0.9871 R2, respectively.