From microarray to biology: an integrated experimental, statistical and in silico analysis of how the extracellular matrix modulates the phenotype of cancer cells.

From microarray to biology: an integrated experimental, statistical and in silico analysis of how the extracellular matrix modulates the phenotype of cancer cells.
复制标题

DOI:
10.1186/1471-2105-9-s9-s4
复制
发表时间:
2008-08-12
期刊:
影响因子:
3
通讯作者:
Hurst RE
Hurst RE
中科院分区:
生物学4区
文献类型:
--
作者:
Dozmorov MG;Kyker KD;Hauser PJ;Saban R;Buethe DD;Dozmorov I;Centola MB;Culkin DJ;Hurst RE

文献摘要

被引文献

相似文献

描述了一种用于分析微阵列数据的统计学稳健和基于生物学的方法,其将独立的生物学知识和数据与全局F-检验相结合,用于寻找感兴趣的基因,当用于假设生成时,其最小化对重复的需要。首先,将每个微阵列归一化为零附近的噪声水平。然后通过鲁棒线性回归对微阵列数据集进行全局调整。第二,通过发现表达比仅表达技术变异性的那些显著更高的变异性的基因来选择捕获对实验条件的显著响应的感兴趣的基因。对表达数据进行聚类并识别包括上游转录调控元件(TREs)、本体和网络或途径在内的感兴趣基因的表达无关性质将数据组织成生物学上有意义的系统。我们证明,当感兴趣的基因的数量是不方便的大,确定一个子集的“信标基因”代表最大的变化将确定生物操纵改变的途径或网络。然后,整个数据集被用来完成由“信标基因”勾勒出的画面。“这允许构建一个系统的结构化模型,可以生成生物学上可检验的假设。我们通过比较在塑料或细胞外基质上培养的细胞来说明这种方法,所述细胞外基质组织来自全基因组转录扫描的2,000多个感兴趣基因的数据集。通过将预测的TREs模式与活性转录因子的实验测定进行比较,证实了所得模型。
A statistically robust and biologically-based approach for analysis of microarray data is described that integrates independent biological knowledge and data with a global F-test for finding genes of interest that minimizes the need for replicates when used for hypothesis generation. First, each microarray is normalized to its noise level around zero. The microarray dataset is then globally adjusted by robust linear regression. Second, genes of interest that capture significant responses to experimental conditions are selected by finding those that express significantly higher variance than those expressing only technical variability. Clustering expression data and identifying expression-independent properties of genes of interest including upstream transcriptional regulatory elements (TREs), ontologies and networks or pathways organizes the data into a biologically meaningful system. We demonstrate that when the number of genes of interest is inconveniently large, identifying a subset of "beacon genes" representing the largest changes will identify pathways or networks altered by biological manipulation. The entire dataset is then used to complete the picture outlined by the "beacon genes." This allow construction of a structured model of a system that can generate biologically testable hypotheses. We illustrate this approach by comparing cells cultured on plastic or an extracellular matrix which organizes a dataset of over 2,000 genes of interest from a genome wide scan of transcription. The resulting model was confirmed by comparing the predicted pattern of TREs with experimental determination of active transcription factors.