Frozen robust multiarray analysis (fRMA)

Frozen robust multiarray analysis (fRMA)
复制标题

DOI:
10.1093/biostatistics/kxp059
复制
发表时间:
2010-04-01
期刊:
影响因子:
2.1
通讯作者:
Irizarry, Rafael A.
Irizarry, Rafael A.
中科院分区:
数学2区
文献类型:
--
作者:
McCall, Matthew N.;Bolstad, Benjamin M.;Irizarry, Rafael A.

文献摘要

被引文献

相似文献

鲁棒多阵列分析(RMA)是最广泛使用的预处理算法的Affyellow和Nimblegen基因表达微阵列。RMA以模块化的方式执行背景校正、标准化和摘要。最后两步需要同时分析多个阵列。跨样本借用信息的能力为RMA提供了各种优势。例如,汇总步骤拟合考虑探针效应的参数模型,假设探针效应在阵列中是固定的,并改进离群值检测。从拟合模型中获得的弹性,允许创建有用的质量度量。然而,对多个阵列的依赖有2个缺点:(1)RMA不能用于必须单独或小批量处理样本的临床环境,以及(2)单独预处理的数据集不具有可比性。我们提出了一种预处理算法,冷冻RMA(fRMA),它允许一个单独或小批量分析微阵列,然后联合收割机的数据进行分析。这是通过利用来自大型公开可用的微阵列数据库的信息来实现的。特别是,探针特定效应和方差的估计值是预先计算和冻结的。然后,使用新的数据集,这些数据集与来自新数组的信息一起使用,以规范化和汇总数据。我们发现,当数据作为单个批次进行分析时,fRMA与RMA相当,并且在分析多个批次时优于RMA。这里描述的方法在R包fRMA中实现,目前可从http://rafalab.jhsph.edu的软件部分下载。
Robust multiarray analysis (RMA) is the most widely used preprocessing algorithm for Affymetrix and Nimblegen gene expression microarrays. RMA performs background correction, normalization, and summarization in a modular way. The last 2 steps require multiple arrays to be analyzed simultaneously. The ability to borrow information across samples provides RMA various advantages. For example, the summarization step fits a parametric model that accounts for probe effects, assumed to be fixed across arrays, and improves outlier detection. Residuals, obtained from the fitted model, permit the creation of useful quality metrics. However, the dependence on multiple arrays has 2 drawbacks: (1) RMA cannot be used in clinical settings where samples must be processed individually or in small batches and (2) data sets preprocessed separately are not comparable. We propose a preprocessing algorithm, frozen RMA (fRMA), which allows one to analyze microarrays individually or in small batches and then combine the data for analysis. This is accomplished by utilizing information from the large publicly available microarray databases. In particular, estimates of probe-specific effects and variances are precomputed and frozen. Then, with new data sets, these are used in concert with information from the new arrays to normalize and summarize the data. We find that fRMA is comparable to RMA when the data are analyzed as a single batch and outperforms RMA when analyzing multiple batches. The methods described here are implemented in the R package fRMA and are currently available for download from the software section of http://rafalab.jhsph.edu.