Optimizer's dilemma: optimization strongly influences model selection in transcriptomic prediction.

Optimizer's dilemma: optimization strongly influences model selection in transcriptomic prediction.
复制标题

DOI:
10.1093/bioadv/vbae004
复制
发表时间:
2024
期刊:
Bioinformatics advances
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

大多数模型可以使用各种优化方法来拟合数据。虽然模型选择在基于机器学习的研究中经常被报道,但优化器并不经常被注意到。我们应用了在Python的scikit-learn包中实现的两种不同的LASSO逻辑回归实现,使用两种不同的优化方法(坐标下降,在liblinear库中实现,以及随机梯度下降,或SGD),从各种泛癌症驱动基因的基因表达中预测突变状态和基因必要性。对于不同的正则化水平,我们比较了优化器之间的性能和模型稀疏性。经过模型选择和调优,我们发现liblinear和SGD倾向于执行自适应。自由线性模型需要对正则化强度进行更广泛的调整,对于高模型稀疏性(更多非零系数)表现最佳,但不需要选择学习率参数。SGD模型需要调整学习率才能表现良好,但随着正则化强度的降低,在不同的模型稀疏度下通常表现得更稳健。考虑到这些权衡,我们认为优化器的选择应该作为模型选择和验证过程的一部分明确报告,以便读者和评审人员更好地理解结果产生的背景。在本研究中用于进行分析的代码可在https://github.com/greenelab/pancancer-evaluation/tree/master/01_stratified_classification上获得。数据集中所有基因的性能/正则化强度曲线可在https://doi.org/10.6084/m9.figshare.22728644上获得。
Most models can be fit to data using various optimization approaches. While model choice is frequently reported in machine-learning-based research, optimizers are not often noted. We applied two different implementations of LASSO logistic regression implemented in Python’s scikit-learn package, using two different optimization approaches (coordinate descent, implemented in the liblinear library, and stochastic gradient descent, or SGD), to predict mutation status and gene essentiality from gene expression across a variety of pan-cancer driver genes. For varying levels of regularization, we compared performance and model sparsity between optimizers. After model selection and tuning, we found that liblinear and SGD tended to perform comparably. liblinear models required more extensive tuning of regularization strength, performing best for high model sparsities (more nonzero coefficients), but did not require selection of a learning rate parameter. SGD models required tuning of the learning rate to perform well, but generally performed more robustly across different model sparsities as regularization strength decreased. Given these tradeoffs, we believe that the choice of optimizers should be clearly reported as a part of the model selection and validation process, to allow readers and reviewers to better understand the context in which results have been generated. The code used to carry out the analyses in this study is available at https://github.com/greenelab/pancancer-evaluation/tree/master/01_stratified_classification. Performance/regularization strength curves for all genes in the dataset are available at https://doi.org/10.6084/m9.figshare.22728644.
DOI: 10.1016/j.celrep.2018.03.046
发表时间: 2018-04-03
期刊: Cell reports
影响因子: 8.8
作者:
Way GP;Sanchez-Vega F;La K;Armenia J;Chatila WK;Luna A;Sander C;Cherniack AD;Mina M;Ciriello G;Schultz N;Cancer Genome Atlas Research Network;Sanchez Y;Greene CS
通讯作者: Greene CS
DOI: 10.1016/s0960-9776(09)70290-5
发表时间: 2009-10-01
期刊: BREAST
影响因子: 3.9
作者:
Albain, Kathy S.;Paik, Soonmyung;van't Veer, Laura
通讯作者: van't Veer, Laura
DOI: 10.1093/bioinformatics/btaa150
发表时间: 2020-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Liu, Renming;Mancuso, Christopher A.;Krishnan, Arjun
通讯作者: Krishnan, Arjun
DOI: 10.1200/jco.2008.18.1370
发表时间: 2009-03-10
影响因子: 45.3
作者:
Parker, Joel S.;Mullins, Michael;Bernard, Philip S.
通讯作者: Bernard, Philip S.
DOI: 10.1126/science.1235122
发表时间: 2013-03-29
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Vogelstein B;Papadopoulos N;Velculescu VE;Zhou S;Diaz LA Jr;Kinzler KW
通讯作者: Kinzler KW