In silico estimates of tissue components in surgical samples based on expression profiling data.

In silico estimates of tissue components in surgical samples based on expression profiling data.
复制标题

DOI:
10.1158/0008-5472.can-10-0021
复制
发表时间:
2010-08-15
期刊:
影响因子:
11.2
通讯作者:
McClelland M
McClelland M
中科院分区:
医学1区
文献类型:
--
作者:
Wang Y;Xia XQ;Jia Z;Sawyers A;Yao H;Wang-Rodriquez J;Mercola D;McClelland M

文献摘要

被引文献

相似文献

许多疾病的组织样本已用于基因表达谱研究,但这些样本所含的细胞类型通常差异很大。这种变化可能会混淆将表达与临床参数相关联的努力。原则上,可以根据分析数据估计每种主要组织成分的比例,并在研究与疾病参数的相关性之前用于对样本进行分类。来自前列腺癌的四个大型基因表达微阵列数据集(其组织成分由病理学家估计)用于测试用于主要组织成分的计算机预测的多元线性回归模型的性能。每个数据集中的十倍交叉验证得出病理学家的预测与计算机预测之间的平均差异,肿瘤成分为 8% 至 14%,基质成分为 13% 至 17%。在使用类似平台和新鲜冷冻样本的独立数据集中,肿瘤的平均差异为 11% 至 12%,基质的平均差异为 12% 至 17%。当这些模型应用于文献中的 219 个“富含肿瘤”样本阵列时,预计几乎四分之一的样本含有 30% 或更少的肿瘤细胞。此外,37 名复发性癌症患者和 42 名非复发性癌症患者之间的平均预测肿瘤含量存在 10.5% 的差异。因此,与组织百分比相关的基因通常也与复发相关。如果不需要这种相关性,则可能会删除一些样本以重新平衡数据集,或者可能会将组织百分比合并到预测算法中。网络服务“CellPred”被设计用于基于表达数据对样本组织成分进行计算机预测。
Tissue samples from many diseases have been used for gene expression profiling studies, but these samples often vary widely in the cell types they contain. Such variation could confound efforts to correlate expression with clinical parameters. In principle, the proportion of each major tissue component can be estimated from the profiling data and used to triage samples before studying correlations with disease parameters. Four large gene expression microarray data sets from prostate cancer, whose tissue components were estimated by pathologists, were used to test the performance of multivariate linear regression models for in silico prediction of major tissue components. Ten-fold cross-validation within each data set yielded average differences between the pathologists' predictions and the in silico predictions of 8% to 14% for the tumor component and 13% to 17% for the stroma component. Across independent data sets that used similar platforms and fresh frozen samples, the average differences were 11% to 12% for tumor and 12% to 17% for stroma. When the models were applied to 219 arrays of “tumor-enriched” samples in the literature, almost one quarter were predicted to have 30% or less tumor cells. Furthermore, there was a 10.5% difference in the average predicted tumor content between 37 recurrent and 42 nonrecurrent cancer patients. As a result, genes that correlated with tissue percentage generally also correlated with recurrence. If such a correlation is not desired, then some samples might be removed to rebalance the data set or tissue percentages might be incorporated into the prediction algorithm. A web service, “CellPred,” has been designed for the in silico prediction of sample tissue components based on expression data.