Structural learning and integrative decomposition of multi-view data

Structural learning and integrative decomposition of multi-view data
复制标题

DOI:
10.1111/biom.13108
复制
发表时间:
2019-09-15
期刊:
影响因子:
1.9
通讯作者:
Li, Gen
Li, Gen
中科院分区:
数学3区
文献类型:
--
作者:
Gaynanova, Irina;Li, Gen

文献摘要

被引文献

相似文献

多视图数据(来自多个来源的相同样本上的数据)的可用性的增加导致了对基于低秩矩阵分解的模型的强烈兴趣。这些模型通过共享和单独的组件来表示每个数据视图,并已成功地应用于探索性降维、视图之间的关联分析和共识聚类。尽管取得了这些进步,但在对部分共享组件建模和确定每种类型(共享/部分共享/单独)的组件数量方面仍然存在挑战。我们制定了一个新的链接组件模型,直接合并部分共享结构。我们将该模型称为SLIDE,用于多视图数据的结构学习和综合分解。与现有的顺序方法相比,所提出的模型拟合和选择技术允许对每种类型的组件数量进行联合识别。在我们的实证研究中,SLIDE在信号估计和组件选择方面都表现出优异的性能。我们进一步说明了癌症基因组图谱库中乳腺癌数据的方法。
The increased availability of multi-view data (data on the same samples from multiple sources) has led to strong interest in models based on low-rank matrix factorizations. These models represent each data view via shared and individual components, and have been successfully applied for exploratory dimension reduction, association analysis between the views, and consensus clustering. Despite these advances, there remain challenges in modeling partially-shared components and identifying the number of components of each type (shared/partially-shared/individual). We formulate a novel linked component model that directly incorporates partially-shared structures. We call this model SLIDE for Structural Learning and Integrative DEcomposition of multi-view data. The proposed model-fitting and selection techniques allow for joint identification of the number of components of each type, in contrast to existing sequential approaches. In our empirical studies, SLIDE demonstrates excellent performance in both signal estimation and component selection. We further illustrate the methodology on the breast cancer data from The Cancer Genome Atlas repository.