Estimating dataset size requirements for classifying DNA microarray data

Estimating dataset size requirements for classifying DNA microarray data
复制标题

DOI:
10.1089/106652703321825928
复制
发表时间:
2003-01-01
影响因子:
1.7
通讯作者:
Mesirov, JP
Mesirov, JP
中科院分区:
生物学4区
文献类型:
--
作者:
Mukherjee, S;Tamayo, P;Mesirov, JP

文献摘要

被引文献

相似文献

介绍了一种利用学习曲线估计微阵列数据分类所需数据集大小的统计方法。我们的目标是使用现有的分类结果来估计未来分类实验的数据集大小要求,并评估使用额外数据构建的分类器的准确性和重要性。该方法是基于拟合逆幂律模型来构建经验学习曲线。它还包括一个排列测试程序,以评估给定数据集大小的分类性能的统计显著性。这一过程适用于几个分子分类问题,代表了广泛的复杂程度。
A statistical methodology for estimating dataset size requirements for classifying microarray data using learning curves is introduced. The goal is to use existing classification results to estimate dataset size requirements for future classification experiments and to evaluate the gain in accuracy and significance of classifiers built with additional data. The method is based on fitting inverse power-law models to construct empirical learning curves. It also includes a permutation test procedure to assess the statistical significance of classification performance for a given dataset size. This procedure is applied to several molecular classification problems representing a broad spectrum of levels of complexity.