Merits of random forests emerge in evaluation of chemometric classifiers by external validation

Merits of random forests emerge in evaluation of chemometric classifiers by external validation
复制标题

DOI:
10.1016/j.aca.2013.09.027
复制
发表时间:
2013-11-01
影响因子:
6.2
通讯作者:
King, R. D.
King, R. D.
中科院分区:
化学1区
文献类型:
--
作者:
Scott, I. M.;Lin, W.;King, R. D.

文献摘要

被引文献

相似文献

现实世界的应用程序将不可避免地需要分歧的样本上的化学计量学分类器进行训练和未知数需要分类。这一点早已被认识到,但缺乏关于哪些分类器在“外部验证”(EV)中表现最好的实证研究,其中未知样本受到相对于用于训练分类器的人群的变化来源的影响。对分析化学中286个分类研究的调查发现,只有6.6%的研究陈述了训练样本和测试样本之间的差异。相反,大多数测试分类器使用来自训练中使用的相同人群的保留或重新验证(通常是交叉验证)。本研究评估了植物和食品材料的NMR和质谱的广泛分类器,来自四个具有不同数据属性的项目(例如,类别的数量和流行程度不同)和分类目标。研究发现,相对于EV,交叉验证在训练集不同来源的样本上的使用是乐观的(例如,不同的基因型、不同的生长条件、不同的作物收获季节)。对于不同任务的分类器评估,我们使用了基于排名的非参数比较和基于排列的显著性检验。虽然潜变量方法(例如,PLSDA)在64%的调查论文中使用,它们是EV中不太成功的分类器之一,正交信号校正适得其反。相反,最好的EV性能是通过处理高维(914-1898个特征)的机器学习方案获得的。随机森林证实了它们对高维的弹性,尽管仅在4.5%的调查论文中使用,但在完整数据上表现最好。大多数其他机器学习分类器都通过特征选择过滤器(ReliefF)进行了改进,但仍然没有超过随机森林。(C)2013爱思唯尔有限公司版权所有。
Real-world applications will inevitably entail divergence between samples on which chemometric classifiers are trained and the unknowns requiring classification. This has long been recognized, but there is a shortage of empirical studies on which classifiers perform best in 'external validation' (EV), where the unknown samples are subject to sources of variation relative to the population used to train the classifier. Survey of 286 classification studies in analytical chemistry found only 6.6% that stated elements of variance between training and test samples. Instead, most tested classifiers using hold-outs or resampling (usually cross-validation) from the same population used in training. The present study evaluated a wide range of classifiers on NMR and mass spectra of plant and food materials, from four projects with different data properties (e.g., different numbers and prevalence of classes) and classification objectives. Use of cross-validation was found to be optimistic relative to EV on samples of different provenance to the training set (e.g., different genotypes, different growth conditions, different seasons of crop harvest). For classifier evaluations across the diverse tasks, we used ranks-based non-parametric comparisons, and permutation-based significance tests. Although latent variable methods (e.g., PLSDA) were used in 64% of the surveyed papers, they were among the less successful classifiers in EV, and orthogonal signal correction was counterproductive. Instead, the best EV performances were obtained with machine learning schemes that coped with the high dimensionality (914-1898 features). Random forests confirmed their resilience to high dimensionality, as best overall performers on the full data, despite being used in only 4.5% of the surveyed papers. Most other machine learning classifiers were improved by a feature selection filter (ReliefF), but still did not out-perform random forests. (C) 2013 Elsevier B.V. All rights reserved.