The removal of multiplicative, systematic bias allows integration of breast cancer gene expression datasets - improving meta-analysis and prediction of prognosis.

The removal of multiplicative, systematic bias allows integration of breast cancer gene expression datasets - improving meta-analysis and prediction of prognosis.
复制标题

DOI:
10.1186/1755-8794-1-42
复制
发表时间:
2008-09-21
影响因子:
2.7
通讯作者:
Clarke, Robert B.
Clarke, Robert B.
中科院分区:
医学3区
文献类型:
--
作者:
Sims, Andrew H.;Smethurst, Graeme J.;Hey, Yvonne;Okoniewski, Michal J.;Pepper, Stuart D.;Howell, Anthony;Miller, Crispin J.;Clarke, Robert B.

文献摘要

参考文献

被引文献

相似文献

在公共领域的基因表达研究的数量正在迅速增加,代表了一个非常宝贵的资源。然而,微阵列特异性偏倚排除了原始转录水平的荟萃分析,即使RNA来自可比较的来源,并已在相同的微阵列平台上使用类似的方案进行处理。在这里,我们证明,使用Affyandroid数据,可以消除大部分这种偏见,允许多个数据集合法地结合起来进行有意义的荟萃分析。生成了一系列比较乳腺癌和正常乳腺细胞系(MCF 7和MCF 10A)的验证数据集,以检查使用不同量的起始RNA、替代方案、不同代的Affytron GeneChip或扫描硬件生成的数据集之间的变异性。我们证明,系统的,乘法的偏见,在RNA,杂交和图像捕获阶段的微阵列实验。发现简单的批平均居中显著降低了实验间变异的水平,从而允许在数据集之间有信心地比较原始转录水平。通过考虑乳腺癌特异性偏倚,我们能够从6项先前发表的研究中收集到迄今为止最大的原发性乳腺肿瘤基因表达数据集(1107)。使用这个元数据集,我们证明,结合更多的数据集或肿瘤导致差异表达基因的更大重叠和更准确的预后预测。然而,这在很大程度上取决于数据集的组成和患者特征。在微阵列实验的许多阶段都会引入倍增的系统性偏倚。当这些被调和时,原始数据可以直接从不同的基因表达数据集整合,从而产生具有增加的统计功效的新的生物学发现。
The number of gene expression studies in the public domain is rapidly increasing, representing a highly valuable resource. However, dataset-specific bias precludes meta-analysis at the raw transcript level, even when the RNA is from comparable sources and has been processed on the same microarray platform using similar protocols. Here, we demonstrate, using Affymetrix data, that much of this bias can be removed, allowing multiple datasets to be legitimately combined for meaningful meta-analyses. A series of validation datasets comparing breast cancer and normal breast cell lines (MCF7 and MCF10A) were generated to examine the variability between datasets generated using different amounts of starting RNA, alternative protocols, different generations of Affymetrix GeneChip or scanning hardware. We demonstrate that systematic, multiplicative biases are introduced at the RNA, hybridization and image-capture stages of a microarray experiment. Simple batch mean-centering was found to significantly reduce the level of inter-experimental variation, allowing raw transcript levels to be compared across datasets with confidence. By accounting for dataset-specific bias, we were able to assemble the largest gene expression dataset of primary breast tumours to-date (1107), from six previously published studies. Using this meta-dataset, we demonstrate that combining greater numbers of datasets or tumours leads to a greater overlap in differentially expressed genes and more accurate prognostic predictions. However, this is highly dependent upon the composition of the datasets and patient characteristics. Multiplicative, systematic biases are introduced at many stages of microarray experiments. When these are reconciled, raw data can be directly integrated from different gene expression datasets leading to new biological findings with increased statistical power.
在412例患者的基于人群的队列中,乳腺癌的固有分子特征。
DOI: 10.1186/bcr1517
发表时间: 2006
影响因子: 7.4
作者:
Calza, Stefano;Hall, Per;Auer, Gert;Bjohle, Judith;Klaar, Sigrid;Kronenwett, Ulrike;T Liu, Edison;Miller, Lance;Ploner, Alexander;Smeds, Johanna;Bergh, Jonas;Pawitan, Yudi
通讯作者: Pawitan, Yudi
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1016/j.cell.2006.09.048
发表时间: 2006-12-01
期刊: CELL
影响因子: 64.5
作者:
Kouros-Mehr, Hosein;Slorach, Euan M.;Werb, Zena
通讯作者: Werb, Zena
DOI: 10.1093/bioinformatics/btg385
发表时间: 2004-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Benito, M;Parker, J;Marron, JS
通讯作者: Marron, JS
DOI: 10.1038/sj.onc.1208561
发表时间: 2005-07-01
期刊: ONCOGENE
影响因子: 8
作者:
Farmer, P;Bonnefoi, H;Iggo, R
通讯作者: Iggo, R