Copula Flows for Synthetic Data Generation

Copula Flows for Synthetic Data Generation
复制标题

用于合成数据生成的 Copula 流

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
M. Deisenroth
M. Deisenroth
中科院分区:
--
文献类型:
--
作者:
Sanket Kamthe;Samuel A. Assefa;M. Deisenroth

文献摘要

被引文献

相似文献

当可用(真实)数据有限或隐私和数据保护标准只允许有限地使用给定数据时,例如在医疗和金融数据集中,生成高保真合成数据的能力至关重要。目前最先进的合成数据生成方法是基于生成模型,如生成对抗网络(gan)。尽管gan在合成数据生成方面取得了显著的成果,但它们往往难以解释。此外,当使用混合实变量和分类变量时,基于gan的方法可能会受到影响。此外,损失函数(判别器损失)设计本身是特定于问题的,也就是说,生成模型可能对它没有明确训练的任务没有用处。在本文中,我们提出使用概率模型作为合成数据生成器。学习数据的概率模型相当于估计数据的密度。基于copula理论,我们将密度估计任务分为两部分,即单变量边缘估计和单变量边缘上的多元copula密度估计。我们使用归一化流来学习联结密度和单变量边际。我们在密度估计和生成高保真合成数据的能力方面对模拟和真实数据集的方法进行了基准测试
The ability to generate high-fidelity synthetic data is crucial when available (real) data is limited or where privacy and data protection standards allow only for limited use of the given data, e.g., in medical and financial data-sets. Current state-of-the-art methods for synthetic data generation are based on generative models, such as Generative Adversarial Networks (GANs). Even though GANs have achieved remarkable results in synthetic data generation, they are often challenging to interpret.Furthermore, GAN-based methods can suffer when used with mixed real and categorical variables.Moreover, loss function (discriminator loss) design itself is problem specific, i.e., the generative model may not be useful for tasks it was not explicitly trained for. In this paper, we propose to use a probabilistic model as a synthetic data generator. Learning the probabilistic model for the data is equivalent to estimating the density of the data. Based on the copula theory, we divide the density estimation task into two parts, i.e., estimating univariate marginals and estimating the multivariate copula density over the univariate marginals. We use normalising flows to learn both the copula density and univariate marginals. We benchmark our method on both simulated and real data-sets in terms of density estimation as well as the ability to generate high-fidelity synthetic data