Mining SOM expression portraits: feature selection and integrating concepts of molecular function.

Mining SOM expression portraits: feature selection and integrating concepts of molecular function.
复制标题

DOI:
10.1186/1756-0381-5-18
复制
发表时间:
2012-10-08
期刊:
影响因子:
4.5
通讯作者:
Binder H
Binder H
中科院分区:
生物学3区
文献类型:
--
作者:
Wirth H;von Bergen M;Binder H

文献摘要

参考文献

被引文献

相似文献

自组织映射(SOM)能够以特定于样本的图像直观地描绘大样本集合的高维数据。对它们的质地的分析提供了所谓的共表达基因的斑点簇,这需要随后的意义过滤和功能解释。我们通过基因排序问题来解决特征选择问题,并使用分子函数的概念来解释所获得的与斑点相关的列表。将基于简单折叠变化度量或基于正则化学生t统计量的不同表达分数应用于斑点相关基因列表,并特别强调微阵列表达数据的误差特征。使用不同的基因集浓缩分析方法对斑点簇进行分析,重点是预定义基因集的过度表达和/或过度表达。所选基因集的后生基因过度表达被映射到SOM图像中,以将基因功能分配给不同的区域。或者,我们使用基因集丰富分数估计了所有研究样本中与集相关的过度表达谱。它还被应用于斑点簇,以生成丰富的基因集列表。我们使用组织身体指数数据集作为说明性示例,该数据集是人体组织表达数据的集合。我们发现,组织相关的点通常包含丰富的基因集群体,与相应组织中的分子过程很好地对应。此外,我们使用SOM数据过滤显示了特殊的内务以及持续弱和高表达的基因集。所提出的方法允许根据簇相关的基因列表和丰富的基因集对SOM转化的表达数据进行全面的下游分析,以用于功能解释。SOM聚类法意味着能够使用选定的SOM点来定义新的基因组,或者验证和/或修改现有的基因组。
Self organizing maps (SOM) enable the straightforward portraying of high-dimensional data of large sample collections in terms of sample-specific images. The analysis of their texture provides so-called spot-clusters of co-expressed genes which require subsequent significance filtering and functional interpretation. We address feature selection in terms of the gene ranking problem and the interpretation of the obtained spot-related lists using concepts of molecular function. Different expression scores based either on simple fold change-measures or on regularized Student’s t-statistics are applied to spot-related gene lists and compared with special emphasis on the error characteristics of microarray expression data. The spot-clusters are analyzed using different methods of gene set enrichment analysis with the focus on overexpression and/or overrepresentation of predefined sets of genes. Metagene-related overrepresentation of selected gene sets was mapped into the SOM images to assign gene function to different regions. Alternatively we estimated set-related overexpression profiles over all samples studied using a gene set enrichment score. It was also applied to the spot-clusters to generate lists of enriched gene sets. We used the tissue body index data set, a collection of expression data of human tissues as an illustrative example. We found that tissue related spots typically contain enriched populations of gene sets well corresponding to molecular processes in the respective tissues. In addition, we display special sets of housekeeping and of consistently weak and high expressed genes using SOM data filtering. The presented methods allow the comprehensive downstream analysis of SOM-transformed expression data in terms of cluster-related gene lists and enriched gene sets for functional interpretation. SOM clustering implies the ability to define either new gene sets using selected SOM spots or to verify and/or to amend existing ones.
DOI: 10.1186/1471-2105-5-125
发表时间: 2004-09-06
期刊: BMC bioinformatics
影响因子: 3
作者:
Aubert J;Bar-Hen A;Daudin JJ;Robin S
通讯作者: Robin S
DOI: 10.1093/bioinformatics/btg307
发表时间: 2003-11-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Eichler, GS;Huang, S;Ingber, DE
通讯作者: Ingber, DE
DOI: 10.1088/1478-3975/7/1/016004
发表时间: 2010-03-01
期刊: PHYSICAL BIOLOGY
影响因子: 2
作者:
Burden, Conrad J.;Binder, Hans
通讯作者: Binder, Hans
DOI: 10.1186/1471-2105-10-47
发表时间: 2009-02-03
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Ackermann, Marit;Strimmer, Korbinian
通讯作者: Strimmer, Korbinian
DOI: 10.1093/nar/gkl435
发表时间: 2006
影响因子: 14.9
作者:
Abdueva D;Skvortsov D;Tavaré S
通讯作者: Tavaré S