Prediction-Oriented Marker Selection (PROMISE): With Application to High-Dimensional Regression.

Prediction-Oriented Marker Selection (PROMISE): With Application to High-Dimensional Regression.
复制标题

DOI:
10.1007/s12561-016-9169-5
复制
发表时间:
2017-06
影响因子:
1
通讯作者:
Lee JJ
Lee JJ
中科院分区:
其他
文献类型:
--
作者:
Kim S;Baladandayuthapani V;Lee JJ

文献摘要

被引文献

相似文献

在个性化医疗中,生物标志物用于根据个体患者的生物标志物/基因组谱来选择成功可能性最高的治疗方法。两个目标是选择能够准确预测治疗结果的重要生物标志物,并剔除不重要的生物标志物以降低生物学和临床验证的成本。由于基因组数据的高维性,这些目标具有挑战性。基于惩罚回归的变量选择方法(例如套索和弹性网)已经产生了有希望的结果。然而,选择合适的惩罚量对于同时实现这两个目标至关重要。基于交叉验证 (CV) 的标准方法通常可以提供高预测精度和高真阳性率,但代价是太多误报。或者,稳定性选择(SS)控制误报的数量,但代价是产生的真阳性太少。为了规避这些问题,我们提出了面向预测的标记选择(PROMISE),它将 SS 与 CV 结合起来,融合了两种方法的优点。我们将 PROMISE 与套索和弹性网络一起应用在数据分析中表明,与 CV 相比,PROMISE 产生稀疏解,误报少,I 型 + II 型误差小,并保持良好的预测精度,真阳性率略有下降。与 SS 相比,PROMISE 提供了更好的预测精度和真阳性率。总之,当目标是最小化误报和最大化预测精度时,PROMISE 可以应用于许多领域来选择正则化参数。
In personalized medicine, biomarkers are used to select therapies with the highest likelihood of success based on an individual patient’s biomarker/genomic profile. Two goals are to choose important biomarkers that accurately predict treatment outcomes and to cull unimportant biomarkers to reduce the cost of biological and clinical verifications. These goals are challenging due to the high dimensionality of genomic data. Variable selection methods based on penalized regression (e.g., the lasso and elastic net) have yielded promising results. However, selecting the right amount of penalization is critical to simultaneously achieving these two goals. Standard approaches based on cross-validation (CV) typically provide high prediction accuracy with high true positive rates but at the cost of too many false positives. Alternatively, stability selection (SS) controls the number of false positives, but at the cost of yielding too few true positives. To circumvent these issues, we propose prediction-oriented marker selection (PROMISE), which combines SS with CV to conflate the advantages of both methods. Our application of PROMISE with the lasso and elastic net in data analysis shows that, compared to CV, PROMISE produces sparse solutions, few false positives, and small type I + type II error, and maintains good prediction accuracy, with a marginal decrease in the true positive rates. Compared to SS, PROMISE offers better prediction accuracy and true positive rates. In summary, PROMISE can be applied in many fields to select regularization parameters when the goals are to minimize false positives and maximize prediction accuracy.