Sample size planning for classification models

Sample size planning for classification models
复制标题

DOI:
10.1016/j.aca.2012.11.007
复制
发表时间:
2013-01-14
影响因子:
6.2
通讯作者:
Popp, Juergen
Popp, Juergen
中科院分区:
化学1区
文献类型:
--
作者:
Beleites, Claudia;Neugebauer, Ute;Popp, Juergen

文献摘要

被引文献

相似文献

在生物光谱学中,用于分类器训练和测试的适当注释和统计独立的样本(例如患者、批次等)稀缺且昂贵。学习曲线将模型性能显示为训练样本大小的函数,并且可以帮助确定训练良好分类器所需的样本大小。然而,建立一个好的模型实际上是不够的:性能还必须得到证明。我们讨论典型的小样本量情况的学习曲线,每类有 5-25 个独立样本。尽管分类模型达到了可接受的性能,但由于同样有限的测试样本量,学习曲线可能会被随机测试的不确定性完全掩盖。因此,我们确定了在验证中实现合理精度所需的测试样本大小,并发现通常需要 75-100 个样本来测试一个好的但不完美的分类器。这样的数据集将允许根据所实现的性能进行精细的样本量规划。我们还演示了如何计算必要的样本量,以显示一个分类器相对于另一个分类器的优越性:这通常需要数百个统计上独立的测试样本,甚至在理论上是不可能的。我们用大约数据集证明了我们的发现。 2550 个单细胞拉曼光谱(五类:红细胞、白细胞和三种肿瘤细胞系 BT-20、MCF-7 和 OCI-AML3)以及广泛的模拟,可以精确确定相关模型的实际性能。 (C) 2012 Elsevier B.V. 保留所有权利。
In biospectroscopy, suitably annotated and statistically independent samples (e.g. patients, batches, etc.) for classifier training and testing are scarce and costly. Learning curves show the model performance as function of the training sample size and can help to determine the sample size needed to train good classifiers. However, building a good model is actually not enough: the performance must also be proven. We discuss learning curves for typical small sample size situations with 5-25 independent samples per class. Although the classification models achieve acceptable performance, the learning curve can be completely masked by the random testing uncertainty due to the equally limited test sample size. In consequence, we determine test sample sizes necessary to achieve reasonable precision in the validation and find that 75-100 samples will usually be needed to test a good but not perfect classifier. Such a data set will then allow refined sample size planning on the basis of the achieved performance. We also demonstrate how to calculate necessary sample sizes in order to show the superiority of one classifier over another: this often requires hundreds of statistically independent test samples or is even theoretically impossible. We demonstrate our findings with a data set of ca. 2550 Raman spectra of single cells (five classes: erythrocytes, leukocytes and three tumour cell lines BT-20, MCF-7 and OCI-AML3) as well as by an extensive simulation that allows precise determination of the actual performance of the models in question. (C) 2012 Elsevier B.V. All rights reserved.