Statistical variation in progressive scrambling

Statistical variation in progressive scrambling
复制标题

DOI:
10.1007/s10822-004-4077-z
复制
发表时间:
2004-07-01
影响因子:
3.5
通讯作者:
Fox, PC
Fox, PC
中科院分区:
生物学3区
文献类型:
--
作者:
Clark, RD;Fox, PC

文献摘要

被引文献

相似文献

最常用于评估偏最小二乘 (PLS) 模型的鲁棒性和预测性的两种方法是交叉验证和响应随机化。然而,这两种方法对于包含冗余观测值的数据集可能过于乐观。在普通最小二乘回归中广泛用于评估模型稳定性的扰动分析仅适用于描述符彼此独立且误差独立且正态分布的情况;这两种假设对于一般的 QSAR 和特别的 PLS 都不成立。渐进置乱是一种新颖的非参数方法,以不干扰数据底层协方差结构的方式扰动响应空间中的模型。在这里,我们引入了对渐进式置乱分析产生的两个特征值的调整 - 已弃用的预测率 (Q(s)(*2)) 和预测标准误差 (SDEPs*) - 纠正引入的扰动的影响。我们还探讨了调整值(Q(0)(*2) 和 SDEP0*)的统计行为以及对扰动的敏感性(dq(2)/dr(yy)(2))。结果表明,这三种统计量对于稳定的 PLS 模型来说都是稳健的,就其确定的随机成分以及由于训练集选择中涉及的采样效应而导致的变化而言。
The two methods most often used to evaluate the robustness and predictivity of partial least squares (PLS) models are cross-validation and response randomization. Both methods may be overly optimistic for data sets that contain redundant observations, however. The kinds of perturbation analysis widely used for evaluating model stability in the context of ordinary least squares regression are only applicable when the descriptors are independent of each other and errors are independent and normally distributed; neither assumption holds for QSAR in general and for PLS in particular. Progressive scrambling is a novel, nonparametric approach to perturbing models in the response space in a way that does not disturb the underlying covariance structure of the data. Here, we introduce adjustments for two of the characteristic values produced by a progressive scrambling analysis - the deprecated predictivity (Q(s)(*2)) and standard error of prediction (SDEPs*) - that correct for the effect of introduced perturbation. We also explore the statistical behavior of the adjusted values (Q(0)(*2) and SDEP0*) and the sensitivity to perturbation (dq(2)/dr(yy)(2)). It is shown that the three statistics are all robust for stable PLS models, in terms of the stochastic component of their determination and of their variation due to sampling effects involved in training set selection.