Cross-platform analysis of cancer microarray data improves gene expression based classification of phenotypes.

Cross-platform analysis of cancer microarray data improves gene expression based classification of phenotypes.
复制标题

DOI:
10.1186/1471-2105-6-265
复制
发表时间:
2005-11-04
期刊:
影响因子:
3
通讯作者:
Brors B
Brors B
中科院分区:
生物学4区
文献类型:
--
作者:
Warnat P;Eils R;Brors B

文献摘要

参考文献

被引文献

相似文献

DNA微阵列技术在细胞转录组表征中的广泛应用导致来自癌症研究的微阵列数据量不断增加。虽然在这些不同的研究中针对相同类型的癌症提出了类似的问题,但由于使用异质性微阵列平台和分析方法,因此对其结果的比较分析受到阻碍。相反,不同的研究结果结合在一个解释性的水平上的荟萃分析方法,我们在这里研究如何直接整合原始的微阵列数据从不同的研究监督分类分析的目的。我们使用中位数排名分数和分位数离散化,从不同的平台获得基因表达的数值可比措施。然后,这些变换后的数据用于基于支持向量机的分类器的训练。我们将这种方法应用于六个公开的癌症微阵列基因表达数据集,这些数据集由三对研究组成,每对研究检查相同类型的癌症,即乳腺癌,前列腺癌或急性髓性白血病。对于每一对,一项研究通过cDNA微阵列进行,另一项通过寡核苷酸微阵列进行。在每一对中,通过对交叉验证分析中从两个数据集中随机选择的数据实例进行训练和测试,实现了高分类准确率(> 85%)。为了验证这种跨平台分类分析的潜力,我们使用了两个白血病微阵列数据集,以表明在综合分析中选择了与白血病生物学相关的重要基因,这些基因在单集分析中均被遗漏。多个癌症微阵列数据集的跨平台分类产生在由不同实验室和微阵列技术生成的大量微阵列样品上发现和验证的有区别的基因表达特征。通过这种方法生成的预测模型比在单个数据集上生成的预测模型更好地验证,同时显示出高的预测能力和改进的泛化性能。
The extensive use of DNA microarray technology in the characterization of the cell transcriptome is leading to an ever increasing amount of microarray data from cancer studies. Although similar questions for the same type of cancer are addressed in these different studies, a comparative analysis of their results is hampered by the use of heterogeneous microarray platforms and analysis methods. In contrast to a meta-analysis approach where results of different studies are combined on an interpretative level, we investigate here how to directly integrate raw microarray data from different studies for the purpose of supervised classification analysis. We use median rank scores and quantile discretization to derive numerically comparable measures of gene expression from different platforms. These transformed data are then used for training of classifiers based on support vector machines. We apply this approach to six publicly available cancer microarray gene expression data sets, which consist of three pairs of studies, each examining the same type of cancer, i.e. breast cancer, prostate cancer or acute myeloid leukemia. For each pair, one study was performed by means of cDNA microarrays and the other by means of oligonucleotide microarrays. In each pair, high classification accuracies (> 85%) were achieved with training and testing on data instances randomly chosen from both data sets in a cross-validation analysis. To exemplify the potential of this cross-platform classification analysis, we use two leukemia microarray data sets to show that important genes with regard to the biology of leukemia are selected in an integrated analysis, which are missed in either single-set analysis. Cross-platform classification of multiple cancer microarray data sets yields discriminative gene expression signatures that are found and validated on a large number of microarray samples, generated by different laboratories and microarray technologies. Predictive models generated by this approach are better validated than those generated on a single data set, while showing high predictive power and improved generalization performance.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1023/a:1016304305535
发表时间: 2002-10-01
影响因子: 4.8
作者:
Liu, H;Hussain, F;Dash, M
通讯作者: Dash, M
DOI: 10.1016/s0140-6736(05)17866-0
发表时间: 2005-02-05
期刊: LANCET
影响因子: 168.9
作者:
Michiels, S;Koscielny, S;Hill, C
通讯作者: Hill, C
DOI: 10.1182/blood.v94.12.4370.424k34_4370_4373
发表时间: 1999-12-15
期刊: BLOOD
影响因子: 20.3
作者:
Cazzaniga, G;Tosi, S;Biondi, A
通讯作者: Biondi, A
DOI: 10.1056/nejmoa031046
发表时间: 2004-04-15
影响因子: 158.5
作者:
Bullinger, L;Döhner, K;Pollack, JR
通讯作者: Pollack, JR