VIGAN: Missing View Imputation with Generative Adversarial Networks.

VIGAN: Missing View Imputation with Generative Adversarial Networks.
复制标题

DOI:
10.1109/bigdata.2017.8257992
复制
发表时间:
2017
期刊:
Proceedings : ... IEEE International Conference on Big Data. IEEE International Conference on Big Data
影响因子:
--
通讯作者:
Bi J
Bi J
中科院分区:
其他
文献类型:
--
作者:
Shang C;Palmer A;Sun J;Chen KS;Lu J;Bi J

文献摘要

被引文献

相似文献

在大数据成为常态的时代,人们对数据的数量关注较少,但更多的是数据的质量和完整性。在许多学科中,数据是从异构源收集的,导致多视图或多模态数据集。在多视图数据分析中,数据缺失问题一直是一个很难解决的问题。特别是,当某些样本丢失整个数据视图时,它会产生丢失视图问题。当无法根据具体信息对此类样本的数据进行估算时,传统的多重估算或矩阵补全方法在此几乎无效。常用的简单方法是删除缺失视图的样本,可以大大减少样本量,从而降低后续分析的统计功效。在本文中,我们提出了一种新的方法,通过生成对抗网络(GAN),我们命名为VIGAN的视图填充。这种方法首先将每个视图视为一个单独的域,并使用来自每个视图的随机采样数据通过GAN识别域到域映射,然后采用多模态去噪自动编码器(DAE)基于跨视图的配对数据从GAN输出重建缺失视图。然后,通过优化GAN和DAE联合,我们的模型使知识集成域映射和视图对应,以有效地恢复丢失的视图。基准数据集上的实证结果验证了VIGAN的方法,通过比较对最先进的。VIGAN在物质使用障碍的遗传研究中的评估进一步证明了这种方法在生命科学中的有效性和可用性。
In an era when big data are becoming the norm, there is less concern with the quantity but more with the quality and completeness of the data. In many disciplines, data are collected from heterogeneous sources, resulting in multi-view or multi-modal datasets. The missing data problem has been challenging to address in multi-view data analysis. Especially, when certain samples miss an entire view of data, it creates the missing view problem. Classic multiple imputations or matrix completion methods are hardly effective here when no information can be based on in the specific view to impute data for such samples. The commonly-used simple method of removing samples with a missing view can dramatically reduce sample size, thus diminishing the statistical power of a subsequent analysis. In this paper, we propose a novel approach for view imputation via generative adversarial networks (GANs), which we name by VIGAN. This approach first treats each view as a separate domain and identifies domain-to-domain mappings via a GAN using randomly-sampled data from each view, and then employs a multi-modal denoising autoencoder (DAE) to reconstruct the missing view from the GAN outputs based on paired data across the views. Then, by optimizing the GAN and DAE jointly, our model enables the knowledge integration for domain mappings and view correspondences to effectively recover the missing view. Empirical results on benchmark datasets validate the VIGAN approach by comparing against the state of the art. The evaluation of VIGAN in a genetic study of substance use disorders further proves the effectiveness and usability of this approach in life science.