Statistical validation of classification and calibration models using bootstrapped Latin partitions

Statistical validation of classification and calibration models using bootstrapped Latin partitions
复制标题

DOI:
10.1016/j.trac.2006.10.010
复制
发表时间:
2006-12-01
影响因子:
13.1
通讯作者:
Harrington, Peter de Boves
Harrington, Peter de Boves
中科院分区:
化学1区
文献类型:
--
作者:
Harrington, Peter de Boves

文献摘要

被引文献

相似文献

对分类和校准方法进行公正的评估很重要,特别是当这些方法适用于日益复杂的数据集时,这些数据集的确定不足。解释任何实验结果都需要精确界,如可信区间。使用自举拉丁划分来评估分类和校准模型,得到了平均预测值的界限。这些界限表征了归因于建立模型和训练集相对于测试集的组成的变异来源。此外,模型可变载荷的平均值上的精度界限允许估计特征特征的重要性。给出了用线性判别分析和模糊规则建立专家系统进行分类的合成数据集,以及用具有一种和三种性质的偏最小二乘回归进行校准的自举拉丁划分的步骤。所有分析都在个人计算机上进行,最长的评估需要6小时的处理时间。方差分析和配对样本t检验也被用来证明这些检验的统计能力。(C)2006年,爱思唯尔有限公司出版。
Unbiased evaluation of classification and calibration methods is important, especially as these methods are applied to increasingly complex data sets that are under-determined. Precision bounds, such as confidence intervals, are required for interpreting any experimental result. Using bootstrapped Latin partitions to evaluate classification and calibration models, bounds on the average predictions were obtained. These bounds characterize sources of variation attributed to building the model and the composition of the training set with respect to the test set. Furthermore, precision bounds on the average of the model-variable loadings allow the significance of characteristic features to be estimated. The procedure for bootstrapped Latin partitions is given and demonstrated with synthetic data sets for classification using linear discriminant analysis and fuzzy rule-building expert systems, and for calibration using partial least squares regression with one and three properties. All analyses were implemented on a personal computer with the longest evaluation requiring 6-h processing time. Analysis of variance and matched sample t-tests were also used to demonstrate the statistical power of these tests. (c) 2006 Published by Elsevier Ltd.