Iterative signature algorithm for the analysis of large-scale gene expression data

Iterative signature algorithm for the analysis of large-scale gene expression data
复制标题

DOI:
10.1103/physreve.67.031902
复制
发表时间:
2003-03-01
期刊:
影响因子:
2.4
通讯作者:
Barkai, N
Barkai, N
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Bergmann, S;Ihmels, J;Barkai, N

文献摘要

被引文献

相似文献

我们提出了一种全基因组表达数据的分析方法。我们的方法旨在克服传统技术的局限性,当应用于大规模数据。而不是分配每个基因到一个单一的集群,我们分配的基因和条件的上下文依赖和潜在的重叠转录模块。我们提供了一个严格的定义的转录模块作为对象从表达数据中检索。建立了一种高效的算法,该算法通过迭代地细化基因和条件的集合来搜索编码在数据中的模块,直到它们与该定义相匹配。每次迭代都涉及由归一化表达式矩阵诱导的线性映射,然后应用阈值函数。我们认为,我们的方法实际上是奇异值分解的推广,这对应于没有阈值的特殊情况。我们分析表明,对于嘈杂的表达数据,我们的方法导致更好的分类,由于实施的阈值。该结果通过基于计算机表达数据的数值分析得到证实。我们简要讨论了通过应用我们的算法从酵母酿酒酵母的表达数据所获得的结果。
We present an approach for the analysis of genome-wide expression data. Our method is designed to overcome the limitations of traditional techniques, when applied to large-scale data. Rather than alloting each gene to a single cluster, we assign both genes and conditions to context-dependent and potentially overlapping transcription modules. We provide a rigorous definition of a transcription module as the object to be retrieved from the expression data. An efficient algorithm, which searches for the modules encoded in the data by iteratively refining sets of genes and conditions until they match this definition, is established. Each iteration involves a linear map, induced by the normalized expression matrix, followed by the application of a threshold function. We argue that our method is in fact a generalization of singular value decomposition, which corresponds to the special case where no threshold is applied. We show analytically that for noisy expression data our approach leads to better classification due to the implementation of the threshold. This result is confirmed by numerical analyses based on in silico expression data. We discuss briefly results obtained by applying our algorithm to expression data from the yeast Saccharomyces cerevisiae.