Unsupervised stratification of cross-validation for accuracy estimation

Unsupervised stratification of cross-validation for accuracy estimation
复制标题

DOI:
10.1016/s0004-3702(99)00094-6
复制
发表时间:
2000-01-01
影响因子:
14.4
通讯作者:
Giakoumakis, EA
Giakoumakis, EA
中科院分区:
计算机科学2区
文献类型:
--
作者:
Diamantidis, NA;Karlis, D;Giakoumakis, EA

文献摘要

被引文献

相似文献

新的学习算法的快速发展增加了对改进的精度估计方法的需求,此外,允许比较几种不同的学习算法的方法对于新的学习算法的性能评估是重要的,在本文中,我们提出了新的精度估计方法,这是k折交叉验证方法的扩展。所提出的方法构建交叉验证折叠确定性,而不是使用随机抽样的方法。折叠的确定性构造使用无监督分层通过利用实例空间中的实例分布来执行。我们的方法是基于一个中心的方法或聚类程序。这些方法试图构建更具代表性的折叠,从而减少所得估计量的偏差。同时,我们的方法允许直接比较学习算法在不同的实验中的性能,因为没有随机性,一个模拟实验研究所提出的方法的性能报告,描绘他们的行为在各种情况下,新的方法主要减少了偏差的估计。(C)2000 Elsevier Science B.V.保留所有权利。
The rapid development of new learning algorithms increases the need for improved accuracy estimation methods, Moreover, methods allowing the comparison of several different learning algorithms are important for the performance evaluation of new ones, In this paper we propose new accuracy estimation methods which are extensions of the k-fold cross-validation method. The methods proposed construct cross-validation folds deterministically instead of using the random sampling approach. The deterministic construction of folds is performed using unsupervised stratification by exploiting the distribution of instances in the instance space. Our methods are based either on the one-center approach or on clustering procedures. These methods attempt to construct more representative folds, therefore reducing the bias of the resulting estimator. At the same time, our methods allow direct comparisons between the performance of learning algorithms in different experiments, since no randomness is present, A simulation experiment examining the performance of the proposed methods is reported, depicting their behavior in a variety of situations, The new methods reduce mainly the bias of the estimator. (C) 2000 Elsevier Science B.V. Ail rights reserved.