Building Normalizing Flows with Stochastic Interpolants

Building Normalizing Flows with Stochastic Interpolants
复制标题

DOI:
10.48550/arxiv.2209.15571
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
M. S. Albergo;E. Vanden-Eijnden
M. S. Albergo;E. Vanden-Eijnden
中科院分区:
其他
文献类型:
--
作者:
M. S. Albergo;E. Vanden-Eijnden

文献摘要

被引文献

相似文献

提出了一种基于任意一对基本概率密度和目标概率密度之间的连续时间归一化流的生成模型。该流的速度场是根据在有限时间内在基础和目标之间插值的时间相关密度的概率电流推断出来的。与基于最大似然原理的传统归一化流推理方法不同,传统的归一化流推理方法需要通过 ODE 求解器进行昂贵的反向传播,而我们的插值方法会导致速度本身产生简单的二次损失,该损失以易于经验估计的期望表示。该流程可用于从基础或目标生成样本,并估计沿插值随时的可能性。此外,还可以优化流程以最小化插值密度的路径长度,从而为构建最佳传输图铺平道路。在基础是高斯密度的情况下,我们还表明,归一化流的速度也可以用于构建扩散模型来对目标进行采样并估计其得分。然而,我们的方法表明,我们可以完全绕过这种扩散,并以更简单的方式在概率流水平上工作,为仅基于常微分方程的方法开辟了一条途径,以替代基于随机微分方程的方法。密度估计任务的基准测试表明,学习流可以以一小部分成本匹配并超越传统的连续流,并且与 CIFAR-10 和 ImageNet 上的图像生成扩散相媲美,成本为 32\times32$。该方法将从头开始的 ODE 流缩放到以前无法达到的图像分辨率,经演示高达 128\times128$。
A generative model based on a continuous-time normalizing flow between any pair of base and target probability densities is proposed. The velocity field of this flow is inferred from the probability current of a time-dependent density that interpolates between the base and the target in finite time. Unlike conventional normalizing flow inference methods based the maximum likelihood principle, which require costly backpropagation through ODE solvers, our interpolant approach leads to a simple quadratic loss for the velocity itself which is expressed in terms of expectations that are readily amenable to empirical estimation. The flow can be used to generate samples from either the base or target, and to estimate the likelihood at any time along the interpolant. In addition, the flow can be optimized to minimize the path length of the interpolant density, thereby paving the way for building optimal transport maps. In situations where the base is a Gaussian density, we also show that the velocity of our normalizing flow can also be used to construct a diffusion model to sample the target as well as estimate its score. However, our approach shows that we can bypass this diffusion completely and work at the level of the probability flow with greater simplicity, opening an avenue for methods based solely on ordinary differential equations as an alternative to those based on stochastic differential equations. Benchmarking on density estimation tasks illustrates that the learned flow can match and surpass conventional continuous flows at a fraction of the cost, and compares well with diffusions on image generation on CIFAR-10 and ImageNet $32\times32$. The method scales ab-initio ODE flows to previously unreachable image resolutions, demonstrated up to $128\times128$.