Predicting sample size required for classification performance.

Predicting sample size required for classification performance.
复制标题

DOI:
10.1186/1472-6947-12-8
复制
发表时间:
2012-02-15
影响因子:
3.5
通讯作者:
Ngo LH
Ngo LH
中科院分区:
医学3区
文献类型:
--
作者:
Figueroa RL;Zeng-Treitler Q;Kandula S;Ngo LH

文献摘要

参考文献

被引文献

相似文献

监督学习方法需要带注释的数据才能生成有效的模型。然而,带注释的数据是一种相对稀缺的资源,并且获取成本可能很高。对于被动和主动学习方法,都需要估计达到性能目标所需的注释样本的大小。我们设计并实现了一种方法,该方法将逆幂律模型拟合到使用小型带注释的训练集创建的给定学习曲线的点。使用非线性加权最小二乘优化进行拟合。然后使用拟合模型来预测分类器的性能和较大样本量的置信区间。为了进行评估,将非线性加权曲线拟合方法应用于使用主动和被动采样方法的临床文本和波形分类任务生成的一组学习曲线,并使用标准拟合优度测量来验证预测。作为对照,我们使用了未加权的拟合方法。总共拟合了 568 个模型,并将模型预测与观察到的性能进行了比较。根据数据集和采样方法,需要 80 到 560 个带注释的样本才能实现平均值和均方根误差低于 0.01。结果还表明,我们的加权拟合方法优于基线未加权方法 (p < 0.05)。本文描述了一种简单有效的样本量预测算法,该算法对学习曲线进行加权拟合。该算法优于先前文献中描述的未加权算法。它可以帮助研究人员确定监督机器学习的注释样本大小。
Supervised learning methods need annotated data in order to generate efficient models. Annotated data, however, is a relatively scarce resource and can be expensive to obtain. For both passive and active learning methods, there is a need to estimate the size of the annotated sample required to reach a performance target. We designed and implemented a method that fits an inverse power law model to points of a given learning curve created using a small annotated training set. Fitting is carried out using nonlinear weighted least squares optimization. The fitted model is then used to predict the classifier's performance and confidence interval for larger sample sizes. For evaluation, the nonlinear weighted curve fitting method was applied to a set of learning curves generated using clinical text and waveform classification tasks with active and passive sampling methods, and predictions were validated using standard goodness of fit measures. As control we used an un-weighted fitting method. A total of 568 models were fitted and the model predictions were compared with the observed performances. Depending on the data set and sampling method, it took between 80 to 560 annotated samples to achieve mean average and root mean squared error below 0.01. Results also show that our weighted fitting method outperformed the baseline un-weighted method (p < 0.05). This paper describes a simple and effective sample size prediction algorithm that conducts weighted fitting of learning curves. The algorithm outperformed an un-weighted algorithm described in previous literature. It can help researchers determine annotation sample size for supervised machine learning.
DOI: 10.1111/1541-0420.00068
发表时间: 2003-09-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
Jiroutek, MR;Muller, KE;Stewart, PW
通讯作者: Stewart, PW
DOI: 10.1109/tpami.2006.156
发表时间: 2006-08-01
影响因子: 23.6
作者:
Li, Mingkun;Sethi, Ishwar K.
通讯作者: Sethi, Ishwar K.
DOI: 10.1198/000313001317098149
发表时间: 2001-08-01
影响因子: 1.8
作者:
Lenth, RV
通讯作者: Lenth, RV
DOI: 10.2307/2531696
发表时间: 1989-09-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
BEAL, SL
通讯作者: BEAL, SL
DOI: 10.1089/106652703321825928
发表时间: 2003-01-01
影响因子: 1.7
作者:
Mukherjee, S;Tamayo, P;Mesirov, JP
通讯作者: Mesirov, JP