A model-based scan statistic for identifying extreme chromosomal regions of gene expression in human tumors

A model-based scan statistic for identifying extreme chromosomal regions of gene expression in human tumors
复制标题

DOI:
10.1093/bioinformatics/bti417
复制
发表时间:
2005-06-15
期刊:
影响因子:
5.8
通讯作者:
Kardia, SLR
Kardia, SLR
中科院分区:
生物学3区
文献类型:
--
作者:
Levin, AM;Ghosh, D;Kardia, SLR

文献摘要

被引文献

相似文献

动机:在染色体背景下分析基因表达数据是癌症研究的最新进展。然而,目前可用的方法无法解释基因,基因密度和基因组特征(例如GC含量)之间的距离的变化,在确定增加或减少染色体区域的基因expression.Results:我们已经开发了一个基于模型的扫描统计,占这些方面的复杂景观的人类基因组中的极端染色体区域的基因表达的识别。为了证明该方法的准确性和实用性,我们将其应用于乳腺癌基因表达数据集,并测试其预测包含中高水平DNA扩增(DNA比率值> 2)的区域的能力。根据扫描统计结果开发了一个分类器,该分类器具有93%的10倍交叉验证分类率和88%的阳性预测值。该结果强烈表明,基于模型的扫描统计和基因表达的增加的染色体区域的表达特征可用于准确地预测含有扩增基因的染色体区域。
Motivation: The analysis of gene expression data in its chromosomal context has been a recent development in cancer research. However, currently available methods fail to account for variation in the distance between genes, gene density and genomic features (e.g. GC content) in identifying increased or decreased chromosomal regions of gene expression.Results: We have developed a model-based scan statistic that accounts for these aspects of the complex landscape of the human genome in the identification of extreme chromosomal regions of gene expression. This method may be applied to gene expression data regardless of the microarray platform used to generate it. To demonstrate the accuracy and utility of this method, we applied it to a breast cancer gene expression dataset and tested its ability to predict regions containing medium-to-high level DNA amplification (DNA ratio values > 2). A classifier was developed from the scan statistic results that had a 10-fold cross-validated classification rate of 93% and a positive predictive value of 88%. This result strongly suggests that the model-based scan statistic and the expression characteristics of an increased chromosomal region of gene expression can be used to accurately predict chromosomal regions containing amplified genes.