Preprocessing Steps for Agilent MicroRNA Arrays: Does the Order Matter?

Preprocessing Steps for Agilent MicroRNA Arrays: Does the Order Matter?
复制标题

DOI:
10.4137/cin.s21630
复制
发表时间:
2014-01-01
期刊:
影响因子:
2
通讯作者:
Zhou, Qin
Zhou, Qin
中科院分区:
其他
文献类型:
--
作者:
Qin, Li-Xuan;Huang, Huei-Chung;Zhou, Qin

文献摘要

被引文献

相似文献

动机/背景:以前关于微阵列预处理的出版物主要集中在单个预处理步骤的方法开发或比较。很少有(如果有的话)专注于推荐预处理步骤的有效排序,特别是与日志转换和探针集汇总相关的标准化。在这项研究中,我们的目的是研究如何相对有序的预处理步骤的影响差异表达分析的Agilent microRNA阵列data.METHODS:一组192未经治疗的原发性妇科肿瘤样本(96子宫内膜肿瘤和96卵巢肿瘤)收集在纪念斯隆-凯特琳癌症中心在2000-2012年期间。从这个相同的样本集,生成两个数据集:一个数据集没有混淆阵列效应的实验设计,并作为基准,另一个数据集表现出阵列效应,并作为测试数据。我们使用以下三个步骤之间的不同顺序预处理我们的测试数据集:分位数归一化,对数转换和中位数汇总。对每个预处理的测试数据集进行差异表达分析,并将结果与基准数据集的结果进行比较。真阳性率、假阳性率和假发现率被用来评估排序的有效性。结果:对数转换、分位数归一化(探针水平数据)和中位数总结的排序略优于其他排序。结论:我们的结果缓解了对排序可能对Agilent microRNA阵列数据分析产生的不确定影响的焦虑。
MOTIVATION/BACKGROUND: Previous publications on microarray preprocessing mostly focused on method development or comparison for an individual preprocessing step. Very few, if any, focused on recommending an effective ordering of the preprocessing steps, in particular, normalization in relationship to log transformation and probe set summarization. In this study, we aim to study how the relative ordering of the preprocessing steps influences differential expression analysis for Agilent microRNA array data.METHODS: A set of 192 untreated primary gynecologic tumor samples (96 endometrial tumors and 96 ovarian tumors) were collected at Memorial Sloan Kettering Cancer Center during the period of 2000-2012. From this same sample set, two datasets were generated: one dataset had no confounding array effects by experimental design and served as the benchmark, and another dataset exhibited array effects and served as the test data. We preprocessed our test dataset using different orderings between the following three steps: quantile normalization, log transformation, and median summarization. Differential expression analysis was performed on each preprocessed test dataset, and the results were compared against the results from the benchmark dataset. True positive rate, false positive rate, and false discovery rate were used to assess the effectiveness of the orderings.RESULTS: The ordering of log transformation, quantile normalization (on probe-level data), and median summarization slightly outperforms the other orderings.CONCLUSION: Our results ease the anxiety over the uncertain effect that the orderings could have on the analysis of Agilent microRNA array data.