How many samples are needed to build a classifier: a general sequential approach

How many samples are needed to build a classifier: a general sequential approach
复制标题

DOI:
10.1093/bioinformatics/bth461
复制
发表时间:
2005-01-01
期刊:
影响因子:
5.8
通讯作者:
Carroll, RJ
Carroll, RJ
中科院分区:
生物学3区
文献类型:
--
作者:
Fu, WJJ;Dougherty, ER;Carroll, RJ

文献摘要

被引文献

相似文献

动机:分类器设计的标准范例是获得特征-标签对的样本,然后应用分类规则从样本数据中导出分类器。通常在实验室情况下,样品大小受到成本、时间或样品材料可用性的限制。因此,调查人员可能希望考虑一个顺序的方法,其中有足够数量的患者训练分类器,以作出合理的决定,同时保持尽可能小的患者数量,使studyoffenders.Results:一个顺序的分类程序研究通过鞅中心极限定理诊断。它在每一步更新分类规则,并提供停止标准,以确保在停止时,未来的受试者将具有小于预定阈值的错误分类概率。模拟研究和应用微阵列数据分析。该方法具有以下几个优点:(1)它顺序地更新分类规则,因此不依赖于其他研究的原始测量值的分布;(2)它在每个顺序步骤评估停止标准,因此可以通过早期停止大大降低成本;以及(3)它不限于任何特定的分类规则,因此适用于任何参数或非参数方法,包括特征选择或提取。
Motivation: The standard paradigm for a classifier design is to obtain a sample of feature-label pairs and then to apply a classification rule to derive a classifier from the sample data. Typically in laboratory situations the sample size is limited by cost, time or availability of sample material. Thus, an investigator may wish to consider a sequential approach in which there is a sufficient number of patients to train a classifier in order to make a sound decision for diagnosis while at the same time keeping the number of patients as small as possible to make the studies affordable.Results: A sequential classification procedure is studied via the martingale central limit theorem. It updates the classification rule at each step and provides stopping criteria to ensure with a certain confidence that at stopping a future subject will have misclassification probability smaller than a predetermined threshold. Simulation studies and applications to microarray data analysis are provided. The procedure possesses several attractive properties: (1) it updates the classification rule sequentially and thus does not rely on distributions of primary measurements from other studies; (2) it assesses the stopping criteria at each sequential step and thus can substantially reduce cost via early stopping; and (3) it is not restricted to any particular classification rule and therefore applies to any parametric or non-parametric method, including feature selection or extraction.