Group Factor Analysis

Group Factor Analysis
复制标题

DOI:
10.1109/tnnls.2014.2376974
复制
发表时间:
2015-09-01
影响因子:
10.4
通讯作者:
Kaski, Samuel
Kaski, Samuel
中科院分区:
计算机科学1区
文献类型:
--
作者:
Klami, Arto;Virtanen, Seppo;Kaski, Samuel

文献摘要

被引文献

相似文献

因子分析(FA)提供描述数据集各个变量之间关系的线性因子。我们将这一经典公式扩展为描述变量组之间关系的线性因子,其中每个组代表一组相关变量或一个数据集。该模型还以一种比以前的扩展更灵活的方式,自然地将典型相关分析扩展到两个以上的集合。我们的解决方案被表述为具有结构稀疏性的潜在变量模型的变分推理,它由两个层次组成:1)较高层次的模型是组之间的关系,2)较低层次的模型是给定较高层次的观察变量。我们表明,所得到的解决方案准确地解决了群因子分析(GFA)问题,优于其他基于fa的解决方案以及更直接的GFA实现。该方法在两个生命科学数据集上进行了演示,一个是关于大脑激活的数据集,另一个是关于系统生物学的数据集,说明了它对不同类型的高维数据源分析的适用性。
Factor analysis (FA) provides linear factors that describe the relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe the relationships between groups of variables, where each group represents either a set of related variables or a data set. The model also naturally extends canonical correlation analysis to more than two sets, in a way that is more flexible than previous extensions. Our solution is formulated as a variational inference of a latent variable model with structural sparsity, and it consists of two hierarchical levels: 1) the higher level models the relationships between the groups and 2) the lower models the observed variables given the higher level. We show that the resulting solution solves the group factor analysis (GFA) problem accurately, outperforming alternative FA-based solutions as well as more straightforward implementations of GFA. The method is demonstrated on two life science data sets, one on brain activation and the other on systems biology, illustrating its applicability to the analysis of different types of high-dimensional data sources.