Subclass mapping: identifying common subtypes in independent disease data sets.

Subclass mapping: identifying common subtypes in independent disease data sets.
复制标题

DOI:
10.1371/journal.pone.0001195
复制
发表时间:
2007-11-21
期刊:
影响因子:
3.7
通讯作者:
Mesirov JP
Mesirov JP
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Hoshida Y;Brunet JP;Tamayo P;Golub TR;Mesirov JP

文献摘要

参考文献

被引文献

相似文献

全基因组表达谱被广泛用于发现疾病的分子亚型。一个仍然存在的挑战是确定在不同平台上产生的多个独立数据集中发现的亚型的对应关系或共性。虽然基于模型的监督学习经常被用于建立这些联系,但模型可能会偏向训练数据集,从而错过测试数据中固有的相关子结构。在此我们描述一种无监督的子类映射方法(SubMap),它揭示独立数据集之间的共同亚型。在应用SubMap之前,一个数据集中的亚型可以通过无监督聚类确定,或者由预先确定的表型给出。我们定义了一种亚型对应性的度量,并基于我们之前在基因集富集分析方面的工作评估其显著性。SubMap方法的优势在于它不会将一个数据集的结构强加于另一个数据集,而是使用一种双向方法来突出两者中的共同子结构。我们展示了这种方法如何揭示几个癌症相关数据集之间的对应关系。值得注意的是,它识别出与雌激素受体状态相关的乳腺癌常见亚型,以及具有相似生存模式的淋巴瘤患者亚组,从而提高了临床结果预测因子的准确性。
Whole genome expression profiles are widely used to discover molecular subtypes of diseases. A remaining challenge is to identify the correspondence or commonality of subtypes found in multiple, independent data sets generated on various platforms. While model-based supervised learning is often used to make these connections, the models can be biased to the training data set and thus miss inherent, relevant substructure in the test data. Here we describe an unsupervised subclass mapping method (SubMap), which reveals common subtypes between independent data sets. The subtypes within a data set can be determined by unsupervised clustering or given by predetermined phenotypes before applying SubMap. We define a measure of correspondence for subtypes and evaluate its significance building on our previous work on gene set enrichment analysis. The strength of the SubMap method is that it does not impose the structure of one data set upon another, but rather uses a bi-directional approach to highlight the common substructures in both. We show how this method can reveal the correspondence between several cancer-related data sets. Notably, it identifies common subtypes of breast cancer associated with estrogen receptor status, and a subgroup of lymphoma patients who share similar survival patterns, thus improving the accuracy of a clinical outcome predictor.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1073/pnas.091062498
发表时间: 2001-04-24
影响因子: 11.1
作者:
Tusher, VG;Tibshirani, R;Chu, G
通讯作者: Chu, G
DOI: 10.1016/s0140-6736(05)17866-0
发表时间: 2005-02-05
期刊: LANCET
影响因子: 168.9
作者:
Michiels, S;Koscielny, S;Hill, C
通讯作者: Hill, C
DOI: 10.1093/jnci/90.15.1138
发表时间: 1998-08-05
影响因子: 10.3
作者:
Lakhani, SR;Jacquemier, J;Easton, DF
通讯作者: Easton, DF
DOI: 10.1038/nmeth757
发表时间: 2005-05-01
期刊: NATURE METHODS
影响因子: 48
作者:
Larkin, JE;Frank, BC;Quackenbush, J
通讯作者: Quackenbush, J