Statistical analysis of high-density oligonucleotide arrays:: a multiplicative noise model

Statistical analysis of high-density oligonucleotide arrays:: a multiplicative noise model
复制标题

DOI:
10.1093/bioinformatics/18.12.1633
复制
发表时间:
2002-12-01
期刊:
影响因子:
5.8
通讯作者:
Corbeil, J
Corbeil, J
中科院分区:
生物学3区
文献类型:
--
作者:
Sásik, R;Calvo, E;Corbeil, J

文献摘要

被引文献

相似文献

动机:高密度寡核苷酸阵列(GeneChip,Affyoung,Santa Clara,CA)已成为许多生物医学研究领域的标准研究工具。它们通过测量来自基因特异性靶标或探针的荧光来同时定量监测数千个基因的表达。信号强度和转录本丰度之间的关系以及标准化问题已经成为最近关注的焦点(Hill等人,2001; Chudin等人,2002; Naef等人,2002年a)。希望研究人员拥有最好的分析工具,以充分利用这种强大的技术所提供的信息。目前有三种可用的分析方法:伴随GeneChip产品的最新发布的Affyssin Microarray Suite 5.0(AMS)软件、Li和Wong的方法(LW; Li和Wong,2001)以及Naef等人的方法(FN; Naef等人,2001年)。AMS方法是为分析单个微阵列而定制的,因此可以用于任何实验设计。另一方面,LW方法依赖于实验中大量的微阵列,并且不能用于分离的微阵列,并且FN方法特定于成对的微阵列,例如由每个“处理”样品具有相应的“对照”样品的实验产生的。我们的重点是分析实验中有一系列的样品。在这种情况下,只能使用AMS、LW和本文所述的方法。本方法是基于模型的,像LW方法,但假设乘性而不是加性噪声,并采用消除统计上显着的离群值,以改善结果。与LW和AMS不同,我们不假设探针特异性背景(通过所谓的失配探针测量)。相反,我们假设均匀的背景,其水平估计使用的失配和完美匹配的探针intensity.Results:我们提出了一种新的方法基因芯片分析,基于统计模型与乘性噪声。我们证明,该方法产生的结果上级优于Affyssoft Microarray Suite 5.0软件和Li和Wong基于模型的方法(Li和Wong,2001)。本方法消除了难以解释的阴性表达指数,并且二元“存在”判定(存在或不存在)被基因表达的统计学显著性(p值)取代。我们已经发现,在(0.1)(16)水平上设定p值阈值产生的“当前”调用数量与AMS软件相同。通过在一对重复的基因芯片(与相同的cRNA杂交)上测试我们的方法,我们发现95.6%的数据点位于1.25倍的区间内。换句话说,我们的方法在1.25倍水平下具有4.4%的I型错误率。LW法的误差率为15%,AMS法的误差率为29%。在本方法中,没有2倍区间之外的点。另一个多次重复实验的方差分析(ANOVA)表明,方差的降低并不伴随着相应的信号降低。相反,本方法的信噪比(由F统计分布测量)平均比AMS好3.4倍,比Li和Wong的好1.4倍。
Motivation: High-density oligonucleotide arrays (GeneChip, Affymetrix, Santa Clara, CA) have become a standard research tool in many areas of biomedical research. They quantitatively monitor the expression of thousands of genes simultaneously by measuring fluorescence from gene-specific targets or probes. The relationship between signal intensities and transcript abundance as well as normalization issues have been the focus of much recent attention (Hill et al., 2001; Chudin et al., 2002; Naef et al., 2002a). It is desirable that a researcher has the best possible analytical tools to make the most of the information that this powerful technology has to offer. At present there are three analytical methods available: the newly released Affymetrix Microarray Suite 5.0 (AMS) software that accompanies the GeneChip product, the method of Li and Wong (LW; Li and Wong, 2001), and the method of Naef et al. (FN; Naef et al., 2001). The AMS method is tailored for analysis of a single microarray, and can therefore be used with any experimental design. The LW method on the other hand depends on a large number of microarrays in an experiment and cannot be used for an isolated microarray, and the FN method is particular to paired microarrays, such as resulting from an experiment in which each 'treatment' sample has a corresponding 'control' sample. Our focus is on analysis of experiments in which there is a series of samples. In this case only the AMS, LW, and the method described in this paper can be used. The present method is model-based, like the LW method, but assumes multiplicative not additive noise, and employs elimination of statistically significant outliers for improved results. Unlike LW and AMS, we do not assume probe-specific background (measured by the so-called mismatch probes). Rather, we assume uniform background, whose level is estimated using both the mismatch and perfect match probe intensities.Results: We present a new method for GeneChip analysis, based on a statistical model with multiplicative noise. We demonstrated that this method yields results superior to those obtained by the Affymetrix Microarray Suite 5.0 software and to those obtained by the model-based method of Li and Wong (Li and Wong, 2001). The present method eliminates the hard-to-interpret negative expression indices, and the binary 'presence' calls (present or absent) are replaced by the statistical significance (p-value) of gene expression. We have found that thresholding the p-values at the (0.1)(16)-level produces about the same number of 'present' calls as the AMS software. By testing our method on a pair of replicate GeneChips (hybridized with the same cRNA), we found that 95.6% of data points lie within the 1.25-fold interval. In other words, our method had a 4.4% type I error rate at the 1.25-fold level. The error rate of the LW method was 15%, and that of the AMS method was 29%. There were no points outside the 2-fold interval with the present method. Analysis of variance (ANOVA) of another experiment with multiple replicates shows that this reduction of variance is not accompanied by a corresponding reduction of signal. On the contrary, the signal-to-noise ratio (as measured by the distribution of F-statistics) of the present method is on average 3.4-times better than that of AMS, and 1.4-times better than that of Li and Wong.