Robust prostate cancer marker genes emerge from direct integration of inter-study microarray data

Robust prostate cancer marker genes emerge from direct integration of inter-study microarray data
复制标题

DOI:
10.1093/bioinformatics/bti647
复制
发表时间:
2005-10-15
期刊:
影响因子:
5.8
通讯作者:
Winslow, RL
Winslow, RL
中科院分区:
生物学3区
文献类型:
--
作者:
Xu, L;Tan, AC;Winslow, RL

文献摘要

被引文献

相似文献

动机:DNA微阵列数据分析以前已被用于识别区分癌症和正常样本的标记基因。然而,由于每个研究的样本量有限,同一癌症的不同研究之间几乎没有共同的标志物。随着微阵列数据的快速积累,整合研究间的微阵列数据以增加样本大小,这可能会导致发现更可靠的markers.Results是非常有趣的:我们提出了一种新的,简单的方法,整合不同的微阵列数据集,以确定标记基因,并将该方法应用于前列腺癌数据集。在这项研究中,通过应用一种新的统计方法,称为最高评分对(TSP)分类器,我们已经确定了一对强大的标记基因(HPN和STAT6),通过整合来自三个不同的前列腺癌研究的微阵列数据集。跨平台验证表明,从标记基因对构建的TSP分类器(其简单地比较相对表达值)在使用各种阵列平台生成的独立数据集上实现了高准确性、灵敏度和特异性。我们的研究结果提出了一种新的模式,从积累的微阵列数据中发现标记基因,并展示了如何利用大量的微阵列数据来提高统计分析的能力。
Motivation: DNA microarray data analysis has been used previously to identify marker genes which discriminate cancer from normal samples. However, due to the limited sample size of each study, there are few common markers among different studies of the same cancer. With the rapid accumulation of microarray data, it is of great interest to integrate inter-study microarray data to increase sample size, which could lead to the discovery of more reliable markers.Results: We present a novel, simple method of integrating different microarray datasets to identify marker genes and apply the method to prostate cancer datasets. In this study, by applying a new statistical method, referred to as the top-scoring pair (TSP) classifier, we have identified a pair of robust marker genes (HPN and STAT6) by integrating microarray datasets from three different prostate cancer studies. Cross-platform validation shows that the TSP classifier built from the marker gene pair, which simply compares relative expression values, achieves high accuracy, sensitivity and specificity on independent datasets generated using various array platforms. Our findings suggest a new model for the discovery of marker genes from accumulated microarray data and demonstrate how the great wealth of microarray data can be exploited to increase the power of statistical analysis.