An empirical study on evaluation metrics of generative adversarial networks

An empirical study on evaluation metrics of generative adversarial networks
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Qiantong Xu;Gao Huang;Yang Yuan;Chuan Guo;Yu Sun;Felix Wu;Kilian Q. Weinberger
Qiantong Xu;Gao Huang;Yang Yuan;Chuan Guo;Yu Sun;Felix Wu;Kilian Q. Weinberger
中科院分区:
其他
文献类型:
--
作者:
Qiantong Xu;Gao Huang;Yang Yuan;Chuan Guo;Yu Sun;Felix Wu;Kilian Q. Weinberger

文献摘要

被引文献

相似文献

评估生成对抗网络(GAN)具有内在的挑战性。在本文中,我们重新审视了几个代表性的基于样本的GAN评估指标,并解决了如何评估评估指标的问题。我们从度量产生有意义的分数的一些必要条件开始,例如区分真实的和生成的样本,识别模式丢弃和模式崩溃,以及检测过拟合。通过一系列精心设计的实验,我们全面研究了现有的基于样本的指标,并确定了它们在实际环境中的优势和局限性。基于这些结果,我们观察到,内核的最大均值离散(MMD)和1-最近邻(1-NN)的两个样本的测试似乎满足大多数理想的属性,只要样本之间的距离计算在一个合适的特征空间。我们的实验还揭示了几种流行的GAN模型行为的有趣特性,例如它们是否记忆训练样本,以及它们距离学习目标分布还有多远。
Evaluating generative adversarial networks (GANs) is inherently challenging. In this paper, we revisit several representative sample-based evaluation metrics for GANs, and address the problem of how to evaluate the evaluation metrics. We start with a few necessary conditions for metrics to produce meaningful scores, such as distinguishing real from generated samples, identifying mode dropping and mode collapsing, and detecting overfitting. With a series of carefully designed experiments, we comprehensively investigate existing sample-based metrics and identify their strengths and limitations in practical settings. Based on these results, we observe that kernel Maximum Mean Discrepancy (MMD) and the 1-Nearest-Neighbor (1-NN) two-sample test seem to satisfy most of the desirable properties, provided that the distances between samples are computed in a suitable feature space. Our experiments also unveil interesting properties about the behavior of several popular GAN models, such as whether they are memorizing training samples, and how far they are from learning the target distribution.