Rank-based methods as a non-parametric alternative of the T-statistic for the analysis of biological microarray data

Rank-based methods as a non-parametric alternative of the T-statistic for the analysis of biological microarray data
复制标题

DOI:
10.1142/s0219720005001442
复制
发表时间:
2005-10-01
影响因子:
1
通讯作者:
Herzyk, Pawel
Herzyk, Pawel
中科院分区:
生物学4区
文献类型:
--
作者:
Breitling, Rainer;Herzyk, Pawel

文献摘要

被引文献

相似文献

最近,我们引入了一个基于秩的检验统计量,秩积(RP),作为一个新的非参数方法检测差异表达基因的微阵列实验。它已被证明可以在生物数据集上产生令人惊讶的良好结果。然而,这种性能的基础和方法的局限性却鲜为人知。在这里,我们探讨了这种基于秩的方法在各种条件下使用模拟的微阵列数据的性能,并将其与经典的Wilcoxon秩和t统计量进行比较,这是大多数替代差异基因表达检测技术的基础。RP对于通过差异表达分选基因比t统计或Wilcoxon秩和更强大和准确-特别是对于低于10的重复数,这是生物实验中最常用的。当数据被非正态随机噪声污染或当样品非常不均匀时,例如因为它们来自不同的时间点或包含受影响和未受影响的细胞的混合物时,其相对性能特别强。然而,RP假设所有基因的测量方差相等,并且当违反此假设时,倾向于给出过于乐观的p值。因此,在计算RP值之前,必须对数据进行适当的方差稳定归一化。如果这是不可能的,RP的另一个基于等级的变体(平均等级)提供了一个有用的替代方案,具有非常相似的整体性能。RP方法的实现可从作者网站(http://www.brc.dcs.gla.ac.uk/glama)下载。
We have recently introduced a rank-based test statistic, RankProducts (RP), as a new non-parametric method for detecting differentially expressed genes in microarray experiments. It has been shown to generate surprisingly good results with biological datasets. The basis for this performance and the limits of the method are, however, little understood. Here we explore the performance of such rank-based approaches under a variety of conditions using simulated microarray data, and compare it with classical Wilcoxon rank sums and t-statistics, which form the basis of most alternative differential gene expression detection techniques.We show that for realistic simulated microarray datasets, RP is more powerful and accurate for sorting genes by differential expression than t-statistics or Wilcoxon rank sums - in particular for replicate numbers below 10, which are most commonly used in biological experiments.Its relative performance is particularly strong when the data are contaminated by non-normal random noise or when the samples are very inhomogenous, e.g. because they come from different time points or contain a mixture of affected and unaffected cells.However, RP assumes equal measurement variance for all genes and tends to give overly optimistic p-values when this assumption is violated. It is therefore essential that proper variance stabilizing normalization is performed on the data before calculating the RP values. Where this is impossible, another rank-based variant of RP (average ranks) provides a useful alternative with very similar overall performance.The Perl scripts implementing the simulation and evaluation are available upon request. Implementations of the RP method are available for download from the authors website (http://www.brc.dcs.gla.ac.uk/glama).