Comments on the analysis of unbalanced microarray data

Comments on the analysis of unbalanced microarray data
复制标题

DOI:
10.1093/bioinformatics/btp363
复制
发表时间:
2009-08-15
期刊:
影响因子:
5.8
通讯作者:
Kerr, Kathleen F.
Kerr, Kathleen F.
中科院分区:
生物学3区
文献类型:
--
作者:
Kerr, Kathleen F.

文献摘要

被引文献

相似文献

动机:排列测试是非常流行的分析微阵列数据,以确定差异表达(DE)的基因,估计错误发现率(FDR)是一个非常流行的方法来解决固有的多重测试问题。然而,结合这些方法可能是有问题的,当样本大小是不平等的。结果:不平衡的数据,排列检验可能是不合适的,因为他们不测试的假设的利益。此外,排列测试可能会有偏差。使用有偏P值估计FDR可能会在这些估计值中产生不可接受的偏差。结果还表明,合并跨基因的置换空分布的方法可能会产生无效的P值,因为即使是非DE基因也可能具有不同的置换空分布。我们鼓励研究人员使用已被证明可以可靠区分DE基因的统计数据,但要注意相关的P值可能无效,或者是区分DE基因的有效性较低的指标。
Motivation: Permutation testing is very popular for analyzing microarray data to identify differentially expressed (DE) genes; estimating false discovery rates (FDRs) is a very popular way to address the inherent multiple testing problem. However, combining these approaches may be problematic when sample sizes are unequal.Results: With unbalanced data, permutation tests may not be suitable because they do not test the hypothesis of interest. In addition, permutation tests can be biased. Using biased P-values to estimate the FDR can produce unacceptable bias in those estimates. Results also show that the approach of pooling permutation null distributions across genes can produce invalid P-values, since even non-DE genes can have different permutation null distributions. We encourage researchers to use statistics that have been shown to reliably discriminate DE genes, but caution that associated P-values may be either invalid, or a less-effective metric for discriminating DE genes.