Independent component analysis reveals new and biologically significant structures in micro array data.

Independent component analysis reveals new and biologically significant structures in micro array data.
复制标题

DOI:
10.1186/1471-2105-7-290
复制
发表时间:
2006-06-08
期刊:
影响因子:
3
通讯作者:
Höglund M
Höglund M
中科院分区:
生物学4区
文献类型:
--
作者:
Frigyesi A;Veerla S;Lindgren D;Höglund M

文献摘要

参考文献

被引文献

相似文献

在微阵列数据中发现有生物学意义的结构的标准方法的另一种选择是将数据视为盲源分离(BSS)问题。BSS试图将混合信号分离成不同的信号源,涉及从几个观察到的线性混合中恢复信号的问题。在微阵列数据的背景下,“来源”可能对应于特定的细胞反应或共调节基因。我们将独立成分分析(ICA)应用于三种不同的微阵列数据集;两个肿瘤数据集和一个时间序列实验。为了获得可靠的成分,我们使用迭代ICA来估计成分的中心型。我们发现许多排名较低的成分确实可能表现出很强的生物学一致性,因此具有生物学意义。总的来说,与基于相关表达的结果相比,ICA获得了更高的分辨率,并且基因本体(GO)类别的基因簇数量更多。此外,鉴定了分子亚型和具有特定染色体易位的肿瘤的特征成分。ICA还发现了多个基因簇对相同的氧化石墨烯类别具有重要意义,因此揭示了更高水平的生物异质性,即使在连贯的基因群中也是如此。尽管ICA方法主要检测隐藏变量,但这些隐藏变量在时间序列数据和肿瘤数据中表现为高度相关的基因。这进一步加强了ICA检测到的潜在变量的生物学相关性。
An alternative to standard approaches to uncover biologically meaningful structures in micro array data is to treat the data as a blind source separation (BSS) problem. BSS attempts to separate a mixture of signals into their different sources and refers to the problem of recovering signals from several observed linear mixtures. In the context of micro array data, "sources" may correspond to specific cellular responses or to co-regulated genes. We applied independent component analysis (ICA) to three different microarray data sets; two tumor data sets and one time series experiment. To obtain reliable components we used iterated ICA to estimate component centrotypes. We found that many of the low ranking components indeed may show a strong biological coherence and hence be of biological significance. Generally ICA achieved a higher resolution when compared with results based on correlated expression and a larger number of gene clusters with significantly enriched for gene ontology (GO) categories. In addition, components characteristic for molecular subtypes and for tumors with specific chromosomal translocations were identified. ICA also identified more than one gene clusters significant for the same GO categories and hence disclosed a higher level of biological heterogeneity, even within coherent groups of genes. Although the ICA approach primarily detects hidden variables, these surfaced as highly correlated genes in time series data and in one instance in the tumor data. This further strengthens the biological relevance of latent variables detected by ICA.
DOI: 10.1186/gb-2003-4-11-r76
发表时间: 2003
期刊: Genome biology
影响因子: 12.3
作者:
Lee SI;Batzoglou S
通讯作者: Batzoglou S
DOI: 10.1038/sj.onc.1207562
发表时间: 2004-08-26
期刊: ONCOGENE
影响因子: 8
作者:
Saidi, SA;Holland, CM;Smith, SK
通讯作者: Smith, SK
DOI: 10.1371/journal.pbio.0020007
发表时间: 2004-02
期刊: PLoS biology
影响因子: 9.8
作者:
Chang HY;Sneddon JB;Alizadeh AA;Sood R;West RB;Montgomery K;Chi JT;van de Rijn M;Botstein D;Brown PO
通讯作者: Brown PO
DOI: 10.1089/cmb.2004.11.1090
发表时间: 2004-12-01
影响因子: 1.7
作者:
Chiappetta, P;Roubaud, MC;Torrésani, B
通讯作者: Torrésani, B
DOI: 10.1056/nejmoa031046
发表时间: 2004-04-15
影响因子: 158.5
作者:
Bullinger, L;Döhner, K;Pollack, JR
通讯作者: Pollack, JR