Exploration, normalization, and summaries of high density oligonucleotide array probe level data

Exploration, normalization, and summaries of high density oligonucleotide array probe level data
复制标题

DOI:
10.1093/biostatistics/4.2.249
复制
发表时间:
2003-04-01
期刊:
影响因子:
2.1
通讯作者:
Speed, TP
Speed, TP
中科院分区:
数学2区
文献类型:
--
作者:
Irizarry, RA;Hobbs, B;Speed, TP

文献摘要

被引文献

相似文献

在本文中,我们报告了从Affymetrix高密度寡核苷酸阵列数据的探索性分析。基因芯片(R)系统,旨在改进目前使用的基因表达测量方法。我们的分析使用了三个数据集:一个由五个MGU74A小鼠基因芯片(R)阵列组成的小型实验研究,部分数据来自Gene Logic和惠氏遗传学研究所进行的一项广泛的研究,涉及95个gh - u95a人类基因芯片(R)阵列;以及Gene Logic进行的稀释研究的一部分,涉及75个HG-U95A基因芯片(R)阵列。我们展示了这些数据的完美匹配和不匹配探针(PM和MM)值的一些常见特征,并检查了来自被认为是有缺陷的探针的探针级数据的方差-均值关系,因此只提供噪声。我们解释了为什么需要使用探针级强度将数组彼此规范化。然后,我们使用峰值数据检查PM和MM的行为,并评估三种常用的汇总测量方法:Affymetrix的(i)平均差(AvDiff)和(ii) MAS 5.0信号,以及(iii) Li和Wong乘法模型基于表达指数(MBEI)。探针级数据的探索性数据分析激发了一种新的汇总度量,即背景调整、归一化和对数转换PM值的鲁棒多阵列平均值(RMA)。我们使用稀释研究数据评估了四种表达汇总方法,从偏倚、方差和(对于MBEI和RMA)模型拟合的角度评估了它们的行为。最后,我们根据使用峰值数据检测已知水平的差异表达的能力来评估算法。我们得出的结论是,使用RMA并使用线性模型将标准误差(SE)附加到该数量上没有明显的缺点,该模型可以消除探针特定的亲和力。包含本文中用于分析的函数的R包是Bioconductor项目的一部分,可以下载(http://www.bioconductor.org)。补充材料,如彩色版本的数字,可在网上(http://www.biostatjhsph.edu/similar toririzarr/affy)。
In this paper we report exploratory analyses of high-density oligonucleotide array data from the Affymetrix. GeneChip(R) system with the objective of improving upon currently used measures of gene expression. Our analyses make use of three data sets: a small experimental study consisting of five MGU74A mouse GeneChip(R) arrays, part of the data from an extensive spike-in study conducted by Gene Logic and Wyeth's Genetics Institute involving 95 HG-U95A human GeneChip(R) arrays; and part of a dilution study conducted by Gene Logic involving 75 HG-U95A GeneChip(R) arrays. We display some familiar features of the perfect match and mismatch probe (PM and MM) values of these data, and examine the variance-mean relationship with probe-level data from probes believed to be defective, and so delivering noise only. We explain why we need to normalize the arrays to one another using probe level intensities. We then examine the behavior of the PM and MM using spike-in data and assess three commonly used summary measures: Affymetrix's (i) average difference (AvDiff) and (ii) MAS 5.0 signal, and (iii) the Li and Wong multiplicative model-based expression index (MBEI). The exploratory data analyses of the probe level data motivate a new summary measure that is a robust multi-array average (RMA) of background-adjusted, normalized, and log-transformed PM values. We evaluate the four expression summary measures using the dilution study data, assessing their behavior in terms of bias, variance and (for MBEI and RMA) model fit. Finally, we evaluate the algorithms in terms of their ability to detect known levels of differential expression using the spike-in data. We conclude that there is no obvious downside to using RMA and attaching a standard error (SE) to this quantity using a linear model which removes probe-specific affinities.An R package with the functions used for the analyses in this paper is part of the Bioconductor project and can be downloaded (http://www.bioconductor.org). Supplemental material, such as color versions of the figures, is available on the web (http://www.biostatjhsph.edu/similar toririzarr/affy).