Discovering biological progression underlying microarray samples.

Discovering biological progression underlying microarray samples.
复制标题

DOI:
10.1371/journal.pcbi.1001123
复制
发表时间:
2011-04
影响因子:
4.3
通讯作者:
Plevritis SK
Plevritis SK
中科院分区:
生物学2区
文献类型:
--
作者:
Qiu P;Gentles AJ;Plevritis SK

文献摘要

参考文献

被引文献

相似文献

在经历分化等过程的生物系统中,存在一个清晰的进展概念。我们提出了一种称为样本进展发现(SPD)的新型计算方法,用于发现微阵列基因表达数据背后的生物进展模式。 SPD 假设微阵列数据集中的各个样本与未知的生物过程(即分化、发育、细胞周期、疾病进展)相关,并且每个样本代表该过程进展中的一个未知点。 SPD 旨在以揭示潜在进展的方式组织样本,并同时识别负责该进展的基因子集。我们在各种微阵列数据集上展示了 SPD 的性能,这些数据集是通过在生物过程进展的不同点进行采样而生成的,而不向 SPD 提供任何基础过程的信息。当应用于细胞周期时间序列微阵列数据集时,SPD 没有提供任何关于样本时间顺序或哪些基因受细胞周期调节的先验知识,但 SPD 恢复了正确的时间顺序并识别了许多与细胞周期相关的基因。当应用于 B 细胞分化数据时,SPD 恢复了正常 B 细胞分化阶段的正确顺序以及 preB-ALL 肿瘤细胞与其细胞来源 preB 之间的联系。当应用于小鼠胚胎干细胞分化数据时,SPD 揭示了 ESC 分化为各种谱系和基因的情况,这些谱系和基因代表了通用过程和谱系特异性过程。当应用于前列腺癌微阵列数据集时,SPD 识别出了反映与疾病阶段一致的进展的基因模块。 SPD 可能最好被视为合成生物学假设的新工具,因为它提供了微阵列数据集基础上可能的生物学进展,也许更重要的是,提供了调节该进展的候选基因。我们提出了一种新颖的计算方法,即样本进展发现(SPD),用于发现微阵列数据集背后的生物进展。与识别样本组之间差异(正常与癌症、治疗与对照)的大多数微阵列数据分析方法相反,SPD 旨在识别样本组内和样本组之间的个体样本之间的潜在进展。我们使用细胞周期、B 细胞分化和小鼠胚胎干细胞分化数据集验证了 SPD 发现生物进展的能力。当应用于进展不清楚的数据集时,我们将 SPD 视为一种假设生成工具。例如,当应用于癌症样本的微阵列数据集时,SPD 假设从个体患者收集的癌症样本代表了癌症发展的内在进展过程中的不同阶段。因此,样本之间的推断关系可能表明癌症进展的轨迹或层次,这可以作为待检验的假设。 SPD不仅限于微阵列数据分析,还可以应用于各种高维数据集。我们使用 MATLAB 图形用户界面实现了 SPD,该界面可从 http://icbp.stanford.edu/software/SPD/ 获取。
In biological systems that undergo processes such as differentiation, a clear concept of progression exists. We present a novel computational approach, called Sample Progression Discovery (SPD), to discover patterns of biological progression underlying microarray gene expression data. SPD assumes that individual samples of a microarray dataset are related by an unknown biological process (i.e., differentiation, development, cell cycle, disease progression), and that each sample represents one unknown point along the progression of that process. SPD aims to organize the samples in a manner that reveals the underlying progression and to simultaneously identify subsets of genes that are responsible for that progression. We demonstrate the performance of SPD on a variety of microarray datasets that were generated by sampling a biological process at different points along its progression, without providing SPD any information of the underlying process. When applied to a cell cycle time series microarray dataset, SPD was not provided any prior knowledge of samples' time order or of which genes are cell-cycle regulated, yet SPD recovered the correct time order and identified many genes that have been associated with the cell cycle. When applied to B-cell differentiation data, SPD recovered the correct order of stages of normal B-cell differentiation and the linkage between preB-ALL tumor cells with their cell origin preB. When applied to mouse embryonic stem cell differentiation data, SPD uncovered a landscape of ESC differentiation into various lineages and genes that represent both generic and lineage specific processes. When applied to a prostate cancer microarray dataset, SPD identified gene modules that reflect a progression consistent with disease stages. SPD may be best viewed as a novel tool for synthesizing biological hypotheses because it provides a likely biological progression underlying a microarray dataset and, perhaps more importantly, the candidate genes that regulate that progression. We present a novel computational approach, Sample Progression Discovery (SPD), to discover biological progression underlying a microarray dataset. In contrast to the majority of microarray data analysis methods which identify differences between sample groups (normal vs. cancer, treated vs. control), SPD aims to identify an underlying progression among individual samples, both within and across sample groups. We validated SPD's ability to discover biological progression using datasets of cell cycle, B-cell differentiation, and mouse embryonic stem cell differentiation. We view SPD as a hypothesis generation tool when applied to datasets where the progression is unclear. For example, when applied to a microarray dataset of cancer samples, SPD assumes that the cancer samples collected from individual patients represent different stages during an intrinsic progression underlying cancer development. The inferred relationship among the samples may therefore indicate a trajectory or hierarchy of cancer progression, which serves as a hypothesis to be tested. SPD is not limited to microarray data analysis, and can be applied to a variety of high-dimensional datasets. We implemented SPD using MATLAB graphical user interface, which is available at http://icbp.stanford.edu/software/SPD/.
前列腺癌的基因表达谱揭示了多个分子途径在转移过程中的参与。
DOI: 10.1186/1471-2407-7-64
发表时间: 2007-04-12
期刊: BMC CANCER
影响因子: 3.8
作者:
Chandran, Uma R.;Ma, Changqing;Dhir, Rajiv;Bisceglia, Michelle;Lyons-Weiler, Maureen;Liang, Wenjing;Michalopoulos, George;Becich, Michael;Monzon, Federico A.
通讯作者: Monzon, Federico A.
DOI: 10.1093/bioinformatics/bti483
发表时间: 2005-07-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Qiu, P;Wang, ZJ;Liu, KJR
通讯作者: Liu, KJR
DOI: 10.1073/pnas.091062498
发表时间: 2001-04-24
影响因子: 11.1
作者:
Tusher, VG;Tibshirani, R;Chu, G
通讯作者: Chu, G
DOI: 10.1016/j.jbi.2007.06.003
发表时间: 2007-12-01
影响因子: 4.5
作者:
Sacchi, Lucia;Larizza, Cristiana;Bellazzi, Riccardo
通讯作者: Bellazzi, Riccardo
DOI: 10.1073/pnas.0504609102
发表时间: 2005-09-06
影响因子: 11.1
作者:
Storey, JD;Xiao, WZ;Davis, RW
通讯作者: Davis, RW