Statistical development and evaluation of microarray gene expression data filters

Statistical development and evaluation of microarray gene expression data filters
复制标题

DOI:
10.1089/cmb.2005.12.482
复制
发表时间:
2005-05-01
影响因子:
1.7
通讯作者:
Cheng, C
Cheng, C
中科院分区:
生物学4区
文献类型:
--
作者:
Pounds, S;Cheng, C

文献摘要

被引文献

相似文献

过滤是一种常见的做法,用于简化微阵列数据的分析,从随后的考虑删除探针集被认为是未表达的。广泛用于Affytek数据分析的m/ n过滤器去除在一组n个芯片中具有少于m个存在调用的所有探针组。m/ n滤波器在应用中没有考虑其统计特性。导出了m/ n滤波器的电平和功率。提出了两种可供选择的滤波器:合并p值滤波器和误差最小化合并p值滤波器.合并p值过滤器将来自存在-不存在p值的信息组合成单个汇总p值,随后将其与选定的显著性阈值进行比较。我们表明,合并的p值过滤器是一个合理的贝塔模型下的统一最强大的统计检验,它表现出更大的权力比m/ n过滤器在模拟研究中考虑的所有情况下。误差最小化合并p值过滤器将汇总p值与阈值进行比较,该阈值基于所有探针的汇总p值的分布分区而确定,以最小化总误差标准.在案例研究分析中,合并p值和误差最小化合并p值滤波器的性能明显优于m/ n滤波器。案例研究分析还证明了一种用于估计通过过滤排除的差异表达探针集的数量以及随后对最终分析的影响的拟议方法。过滤器影响分析表明,即使是最好的过滤器的使用可能会阻碍,而不是提高,发现感兴趣的探针集或基因的能力。实现合并p值和误差最小化合并p值滤波器的S- plus和R例程已经开发出来,可从www. stjudesearch. org/ depts/ biostats/ index. HTML.
Filtering is a common practice used to simplify the analysis of microarray data by removing from subsequent consideration probe sets believed to be unexpressed. The m/ n filter, which is widely used in the analysis of Affymetrix data, removes all probe sets having fewer than m present calls among a set of n chips. The m/ n filter has been widely used without considering its statistical properties. The level and power of the m/ n filter are derived. Two alternative filters, the pooled p- value filter and the error- minimizing pooled p- value filter are proposed. The pooled p- value filter combines information from the present - absent p- values into a single summary p- value which is subsequently compared to a selected significance threshold. We show that the pooled p- value filter is the uniformly most powerful statistical test under a reasonable beta model and that it exhibits greater power than the m/ n filter in all scenarios considered in a simulation study. The error- minimizing pooled p- value filter compares the summary p- value with a threshold determined to minimize a total- error criterion based on a partition of the distribution of all probes' summary p- values. The pooled p- value and error- minimizing pooled p- value filters clearly perform better than the m/ n filter in a case- study analysis. The case- study analysis also demonstrates a proposed method for estimating the number of differentially expressed probe sets excluded by filtering and subsequent impact on the final analysis. The filter impact analysis shows that the use of even the best filter may hinder, rather than enhance, the ability to discover interesting probe sets or genes. S- plus and R routines to implement the pooled p- value and error- minimizing pooled p- value filters have been developed and are available from www. stjuderesearch. org/ depts/ biostats/ index. html.