Error, reproducibility and sensitivity: a pipeline for data processing of Agilent oligonucleotide expression arrays

Error, reproducibility and sensitivity: a pipeline for data processing of Agilent oligonucleotide expression arrays
复制标题

DOI:
10.1186/1471-2105-11-344
复制
发表时间:
2010-06-24
期刊:
影响因子:
3
通讯作者:
Noursadeghi, Mahdad
Noursadeghi, Mahdad
中科院分区:
生物学4区
文献类型:
--
作者:
Chain, Benjamin;Bowen, Helen;Noursadeghi, Mahdad

文献摘要

被引文献

相似文献

背景:表达微阵列越来越多地用于获得大范围生物样品的大规模转录组学信息。尽管如此,对于处理数据、设计实验和分析结果的最佳方式仍有很多争论。此外,文献中许多更复杂的数据分析数学方法仍然无法进入生物研究界。在这项研究中,我们研究的方法提取和分析使用安捷伦长寡核苷酸转录组学平台获得的一个大的数据集,应用于一组人类巨噬细胞和树突状细胞samples.Results:我们描述和验证了一系列的数据提取,转换和归一化步骤,通过一个新的R函数实现。重复标准化参考数据的分析表明,阵列内变异性很小(仅为平均对数信号的2%左右),而重复阵列测量的阵列间变异性的标准差(SD)约为0.5 log(2)单位(平均值的6%)。利用Cy 5/Cy 3信号的比率工作的常见实践在减少误差方面提供了很少的进一步改进。与使用拟南芥样品获得的表达数据的比较表明,每个样品中显示低水平转录的大量基因反映了细胞转录组的真实的复杂性。多维缩放用于显示经处理的数据识别反映定义数据集的一些关键生物变量的底层结构。这种结构是稳健的,可以对多年来收集的样本和由各种操作员收集的样本进行可靠的比较。结论:本研究概述了一个稳健且易于实施的管道,用于从安捷伦表达平台提取、转换、标准化和可视化转录组阵列数据。该分析用于获得由实验(非生物学)阵列内和阵列间变异性引起的SD的定量估计,以及用于确定单个基因是否表达的较低阈值。该研究为真核细胞系统生物学的进一步深入研究提供了可靠的基础。
Background: Expression microarrays are increasingly used to obtain large scale transcriptomic information on a wide range of biological samples. Nevertheless, there is still much debate on the best ways to process data, to design experiments and analyse the output. Furthermore, many of the more sophisticated mathematical approaches to data analysis in the literature remain inaccessible to much of the biological research community. In this study we examine ways of extracting and analysing a large data set obtained using the Agilent long oligonucleotide transcriptomics platform, applied to a set of human macrophage and dendritic cell samples.Results: We describe and validate a series of data extraction, transformation and normalisation steps which are implemented via a new R function. Analysis of replicate normalised reference data demonstrate that intrarray variability is small (only around 2% of the mean log signal), while interarray variability from replicate array measurements has a standard deviation (SD) of around 0.5 log(2) units (6% of mean). The common practise of working with ratios of Cy5/Cy3 signal offers little further improvement in terms of reducing error. Comparison to expression data obtained using Arabidopsis samples demonstrates that the large number of genes in each sample showing a low level of transcription reflect the real complexity of the cellular transcriptome. Multidimensional scaling is used to show that the processed data identifies an underlying structure which reflect some of the key biological variables which define the data set. This structure is robust, allowing reliable comparison of samples collected over a number of years and collected by a variety of operators.Conclusions: This study outlines a robust and easily implemented pipeline for extracting, transforming normalising and visualising transcriptomic array data from Agilent expression platform. The analysis is used to obtain quantitative estimates of the SD arising from experimental (non biological) intra- and interarray variability, and for a lower threshold for determining whether an individual gene is expressed. The study provides a reliable basis for further more extensive studies of the systems biology of eukaryotic cells.