A greedy regression algorithm with coarse weights offers novel advantages.

A greedy regression algorithm with coarse weights offers novel advantages.
复制标题

DOI:
10.1038/s41598-022-09415-2
复制
发表时间:
2022-03-31
期刊:
影响因子:
4.6
通讯作者:
Wilhelmsen KC
Wilhelmsen KC
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Jeffries CD;Ford JR;Tilson JL;Perkins DO;Bost DM;Filer DL;Wilhelmsen KC

文献摘要

参考文献

相似文献

正则化回归分析是一种成熟的分析方法,用于确定预测结果的变量的加权和。我们提出了一种新的粗近似线性函数(CALF)节俭地选择重要的预测因子,并建立简单但功能强大的预测模型。CALF是一种线性回归策略,应用于使用非零权重+1或-1的归一化数据。待优化的定性(线性不变)度量可以是(对于二进制响应)Welch(Student)t检验p值或受试者操作特征的曲线下面积(AUC),或(对于真实的响应)Pearson相关性。在开发风险预测模型时,预测权重至关重要。虽然违反直觉,但事实上,定性指标可以支持具有± 1个权重的CALF,而不是产生真实的数字权重的算法。此外,虽然回归方法可以被预期在输入数据(例如,丢弃数百个的单个主题)CALF权重通常不这样改变。类似地,应用于共线或近似共线变量的一些回归方法产生不可预测的权重大小或方向(在p空间中)作为向量。相比之下,对于CALF,如果一些预测变量是线性相关的或接近线性相关的,CALF只选择最多一个(最具信息性的,如果有的话),忽略其他变量,从而避免在模型中包含两个或更多的共线性变量。
Regularized regression analysis is a mature analytic approach to identify weighted sums of variables predicting outcomes. We present a novel Coarse Approximation Linear Function (CALF) to frugally select important predictors and build simple but powerful predictive models. CALF is a linear regression strategy applied to normalized data that uses nonzero weights + 1 or − 1. Qualitative (linearly invariant) metrics to be optimized can be (for binary response) Welch (Student) t-test p-value or area under curve (AUC) of receiver operating characteristic, or (for real response) Pearson correlation. Predictor weighting is critically important when developing risk prediction models. While counterintuitive, it is a fact that qualitative metrics can favor CALF with ± 1 weights over algorithms producing real number weights. Moreover, while regression methods may be expected to change most or all weight values upon even small changes in input data (e.g., discarding a single subject of hundreds) CALF weights generally do not so change. Similarly, some regression methods applied to collinear or nearly collinear variables yield unpredictable magnitude or the direction (in p-space) of the weights as a vector. In contrast, with CALF if some predictors are linearly dependent or nearly so, CALF simply chooses at most one (the most informative, if any) and ignores the others, thus avoiding the inclusion of two or more collinear variables in the model.
DOI: 10.1016/j.schres.2015.09.008
发表时间: 2015-12
影响因子: 4.5
作者:
Perkins DO;Jeffries CD;Cornblatt BA;Woods SW;Addington J;Bearden CE;Cadenhead KS;Cannon TD;Heinssen R;Mathalon DH;Seidman LJ;Tsuang MT;Walker EF;McGlashan TH
通讯作者: McGlashan TH
DOI: 10.1093/schbul/sbp027
发表时间: 2009-09-01
影响因子: 6.6
作者:
Woods, Scott W.;Addington, Jean;McGlashan, Thomas H.
通讯作者: McGlashan, Thomas H.
DOI: 10.1016/j.scog.2017.10.001
发表时间: 2018-03
期刊: Schizophrenia research. Cognition
影响因子: --
作者:
Ramsay IS;Ma S;Fisher M;Loewy RL;Ragland JD;Niendam T;Carter CS;Vinogradov S
通讯作者: Vinogradov S
DOI: 10.1371/journal.pone.0175683
发表时间: 2017
期刊: PloS one
影响因子: 3.7
作者:
Salvador R;Radua J;Canales-Rodríguez EJ;Solanes A;Sarró S;Goikolea JM;Valiente A;Monté GC;Natividad MDC;Guerrero-Pedraza A;Moro N;Fernández-Corcuera P;Amann BL;Maristany T;Vieta E;McKenna PJ;Pomarol-Clotet E
通讯作者: Pomarol-Clotet E
DOI: 10.1016/j.bbalip.2013.01.002
发表时间: 2013-04
期刊: Biochimica et biophysica acta
影响因子: --
作者:
Lord CC;Thomas G;Brown JM
通讯作者: Brown JM