Sample size and statistical power considerations in high-dimensionality data settings: a comparative study of classification algorithms.

Sample size and statistical power considerations in high-dimensionality data settings: a comparative study of classification algorithms.
复制标题

DOI:
10.1186/1471-2105-11-447
复制
发表时间:
2010-09-03
期刊:
影响因子:
3
通讯作者:
Balasubramanian R
Balasubramanian R
中科院分区:
生物学4区
文献类型:
--
作者:
Guo Y;Graber A;McBurney RN;Balasubramanian R

文献摘要

参考文献

被引文献

相似文献

使用“组学”技术生成的数据具有高维性的特点,每个受试者测量的特征数量大大超过了研究中的受试者数量。在本文中,我们考虑相关的问题,在生物医学研究的设计,其目标是发现的一个子集的功能和相关的算法,可以预测一个二进制的结果,如疾病状态。我们比较了四种常用分类器(K-最近邻,微阵列预测分析,随机森林和支持向量机)在高维数据设置中的性能。我们评估了不同水平的信号噪声比的数据集,类分布的不平衡和选择的量化性能的分类度量的影响。为了指导研究设计,我们提出了一个总结的几个人类或动物模型实验中的“组学”数据的关键特征,利用高含量的质谱和多重免疫分析为基础的技术。对七项“组学”研究数据的分析表明,与动物研究相比,在人类研究中观察到的效应量的平均幅度明显较低。在人体研究中测量的数据的特点是较高的生物变异和离群值的存在。仿真研究结果表明,当类条件特征分布为高斯分布且结果分布为平衡分布时,微阵列分类器预测分析(PAM)具有最高的功效。随机森林是最佳的特征分布时,偏斜和类分布时,不平衡。我们提供了一个免费的开源R统计软件库(MVpower),实现了本文提出的模拟策略。没有一个分类器在所有设置下都具有最佳性能。模拟研究为涉及高维数据的生物医学研究的设计提供了有用的指导。
Data generated using 'omics' technologies are characterized by high dimensionality, where the number of features measured per subject vastly exceeds the number of subjects in the study. In this paper, we consider issues relevant in the design of biomedical studies in which the goal is the discovery of a subset of features and an associated algorithm that can predict a binary outcome, such as disease status. We compare the performance of four commonly used classifiers (K-Nearest Neighbors, Prediction Analysis for Microarrays, Random Forests and Support Vector Machines) in high-dimensionality data settings. We evaluate the effects of varying levels of signal-to-noise ratio in the dataset, imbalance in class distribution and choice of metric for quantifying performance of the classifier. To guide study design, we present a summary of the key characteristics of 'omics' data profiled in several human or animal model experiments utilizing high-content mass spectrometry and multiplexed immunoassay based techniques. The analysis of data from seven 'omics' studies revealed that the average magnitude of effect size observed in human studies was markedly lower when compared to that in animal studies. The data measured in human studies were characterized by higher biological variation and the presence of outliers. The results from simulation studies indicated that the classifier Prediction Analysis for Microarrays (PAM) had the highest power when the class conditional feature distributions were Gaussian and outcome distributions were balanced. Random Forests was optimal when feature distributions were skewed and when class distributions were unbalanced. We provide a free open-source R statistical software library (MVpower) that implements the simulation strategy proposed in this paper. No single classifier had optimal performance under all settings. Simulation studies provide useful guidance for the design of biomedical studies involving high-dimensionality data.
DOI: 10.1111/1467-9868.00346
发表时间: 2002-01-01
影响因子: 5.8
作者:
Storey, JD
通讯作者: Storey, JD
DOI: 10.1093/bioinformatics/bti448
发表时间: 2005-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Pawitan, Y;Michiels, S;Ploner, A
通讯作者: Ploner, A
DOI: 10.1093/biostatistics/kxh015
发表时间: 2005-01-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Dobbin, K;Simon, R
通讯作者: Simon, R
DOI: 10.1198/016214504000001646
发表时间: 2004-12-01
影响因子: 3.7
作者:
M端ller, P;Parmigiani, G;Rousseau, J
通讯作者: Rousseau, J
DOI: 10.1073/pnas.0601231103
发表时间: 2006-04-11
影响因子: 11.1
作者:
Ein-Dor, L;Zuk, O;Domany, E
通讯作者: Domany, E