Blind source separation and the analysis of microarray data

Blind source separation and the analysis of microarray data
复制标题

DOI:
10.1089/cmb.2004.11.1090
复制
发表时间:
2004-12-01
影响因子:
1.7
通讯作者:
Torrésani, B
Torrésani, B
中科院分区:
生物学4区
文献类型:
--
作者:
Chiappetta, P;Roubaud, MC;Torrésani, B

文献摘要

被引文献

相似文献

我们开发了一种基于盲源分离技术的基因表达数据探索性分析方法。这种方法利用高阶统计来识别表达谱的线性模型,描述为“独立来源”的线性组合。因此,它产生了“基本表达模式”(“来源”),这可能被解释为潜在的调控途径。对所获得的来源的进一步分析表明,它们通常以少量特异性共表达或抗表达基因为特征。此外,表达谱到估计来源上的投影通常提供显著的病症聚类。该算法依赖于随机初始化的“独立成分分析”的大量运行,然后搜索“共识源”。“然后,它提供了独立来源的估计,以及对其稳健性的评估。两个数据集(即乳腺癌数据和枯草芽孢杆菌硫代谢数据)上获得的结果表明,一些获得的基因家族对应于众所周知的共调节基因家族,这验证了所提出的方法。
We develop an approach for the exploratory analysis of gene expression data, based upon blind source separation techniques. This approach exploits higher-order statistics to identify a linear model for ( logarithms of) expression profiles, described as linear combinations of "independent sources." As a result, it yields "elementary expression patterns" ( the "sources"), which may be interpreted as potential regulation pathways. Further analysis of the so-obtained sources show that they are generally characterized by a small number of specific coexpressed or antiexpressed genes. In addition, the projections of the expression profiles onto the estimated sources often provides significant clustering of conditions. The algorithm relies on a large number of runs of "independent component analysis" with random initializations, followed by a search of "consensus sources." It then provides estimates for independent sources, together with an assessment of their robustness. The results obtained on two datasets ( namely, breast cancer data and Bacillus subtilis sulfur metabolism data) show that some of the obtained gene families correspond to well known families of coregulated genes, which validates the proposed approach.