oFVSD: A Python package of optimized forward variable selection decoder for high-dimensional neuroimaging data

oFVSD: A Python package of optimized forward variable selection decoder for high-dimensional neuroimaging data
复制标题

oFVSD:用于高维神经影像数据的优化前向变量选择解码器的 Python 包

DOI:
10.1101/2022.12.25.521906
复制
发表时间:
2022
期刊:
bioRxiv
影响因子:
--
通讯作者:
Machizawa Maro G.
Machizawa Maro G.
中科院分区:
--
文献类型:
--
作者:
Dang Tung;Fermin Alan S. R.;Machizawa Maro G.

文献摘要

相似文献

神经成像数据的复杂性和高维度给机器学习(ML)模型解码信息带来了问题,因为特征的数量通常比观察的数量大得多。特征选择是在解码中确定有意义的目标特征的关键步骤之一;然而,使用传统的ML模型,从这种高维神经成像数据中优化特征选择一直是一个挑战。在这里,我们介绍了一个高效和高性能的解码包,其中包含前向变量选择(FVS)算法和超参数优化,可以自动识别分类和回归模型的最佳特征对,默认情况下共实现了18个ML模型。首先,FVS算法使用k折交叉验证步骤评估不同模型的拟合优度,该步骤基于每个模型的预定义标准识别最佳特征子集。接下来,每个ML模型的超参数在每次前向迭代中被优化。最终输出突出显示了每个模型的优化数量的选定特征(感兴趣的大脑区域)及其准确性。此外,该工具箱可以在并行环境中执行,以便在典型的个人计算机上进行有效的计算。通过优化的前向变量选择解码器(oFVSD)流水线,我们在1,113个结构磁共振成像(MRI)数据集上验证了解码性别分类和年龄范围回归的有效性。与没有FVS算法的ML模型和使用Boruta算法作为变量选择对应物的ML模型相比,我们证明了oFVSD在所有ML模型中的表现明显优于没有FVS的对应模型(相关系数r增加约0.20,使用回归模型,分类模型平均增加8%)和Boruta变量选择算法(回归中约0.07%的改进和分类模型中约4%的改进)。此外,我们证实了并行计算的使用大大降低了计算负担的高维MRI数据。总而言之,oFVSD工具箱有效地提高了分类和回归ML模型的性能,提供了MRI数据集的用例示例。由于其灵活性,oFVSD在神经影像学中具有许多其他模式的潜力。这个开源和免费提供的Python包使其成为寻求提高解码准确性的研究社区的宝贵工具箱。
The complexity and high dimensionality of neuroimaging data pose problems for decoding information with machine learning (ML) models because the number of features is often much larger than the number of observations. Feature selection is one of the crucial steps for determining meaningful target features in decoding; however, optimizing the feature selection from such high-dimensional neuroimaging data has been challenging using conventional ML models. Here, we introduce an efficient and high-performance decoding package incorporating a forward variable selection (FVS) algorithm and hyper-parameter optimization that automatically identifies the best feature pairs for both classification and regression models, where a total of 18 ML models are implemented by default. First, the FVS algorithm evaluates the goodness-of-fit across different models using the k-fold cross-validation step that identifies the best subset of features based on a predefined criterion for each model. Next, the hyperparameters of each ML model are optimized at each forward iteration. Final outputs highlight an optimized number of selected features (brain regions of interest) for each model with its accuracy. Furthermore, the toolbox can be executed in a parallel environment for efficient computation on a typical personal computer. With the optimized forward variable selection decoder (oFVSD) pipeline, we verified the effectiveness of decoding sex classification and age range regression on 1,113 structural magnetic resonance imaging (MRI) datasets. Compared to ML models without the FVS algorithm and with the Boruta algorithm as a variable selection counterpart, we demonstrate that the oFVSD significantly outperformed across all of the ML models over the counterpart models without FVS (approximately 0.20 increase in correlation coefficient,r, with regression models and 8% increase in classification models on average) and with Boruta variable selection algorithm (approximately 0.07 improvement in regression and 4% in classification models). Furthermore, we confirmed the use of parallel computation considerably reduced the computational burden for the high-dimensional MRI data. Altogether, the oFVSD toolbox efficiently and effectively improves the performance of both classification and regression ML models, providing a use case example on MRI datasets. With its flexibility, oFVSD has the potential for many other modalities in neuroimaging. This open-source and freely available Python package makes it a valuable toolbox for research communities seeking improved decoding accuracy.