Understanding the Distillation Process from Deep Generative Models to Tractable Probabilistic Circuits

Understanding the Distillation Process from Deep Generative Models to Tractable Probabilistic Circuits
复制标题

DOI:
10.48550/arxiv.2302.08086
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Xuejie Liu;Anji Liu;Guy Van den Broeck;Yitao Liang
Xuejie Liu;Anji Liu;Guy Van den Broeck;Yitao Liang
中科院分区:
其他
文献类型:
--
作者:
Xuejie Liu;Anji Liu;Guy Van den Broeck;Yitao Liang

文献摘要

相似文献

概率电路(PC)是一个通用的、统一的计算框架,用于支持各种推理任务的有效计算(例如,计算边际概率)的易处理的概率模型。为了在复杂的现实世界任务中实现这样的推理能力,Liu等人。(2022)提出从不太容易处理但更具表现力的深层生成模型中提取知识(通过潜在变量赋值)。然而,目前还不清楚是什么因素让这种蒸馏工作得很好。在本文中,我们从理论和经验上发现,PC的性能可以超过它的教师模型。因此,我们不是从最具表现力的深层生成模型进行精馏,而是研究教师模型和PC应该具有什么属性才能获得良好的精馏性能。这导致了通用算法的改进,以及现有潜变量蒸馏管道上的其他特定数据类型的改进。根据经验,在具有挑战性的图像建模基准方面,我们的表现远远超过SOTA TPM。特别是,在ImageNet32上,PC实现了每维4.06比特,仅比变分扩散模型低0.34(Kingma等人,2021)。
Probabilistic Circuits (PCs) are a general and unified computational framework for tractable probabilistic models that support efficient computation of various inference tasks (e.g., computing marginal probabilities). Towards enabling such reasoning capabilities in complex real-world tasks, Liu et al. (2022) propose to distill knowledge (through latent variable assignments) from less tractable but more expressive deep generative models. However, it is still unclear what factors make this distillation work well. In this paper, we theoretically and empirically discover that the performance of a PC can exceed that of its teacher model. Therefore, instead of performing distillation from the most expressive deep generative model, we study what properties the teacher model and the PC should have in order to achieve good distillation performance. This leads to a generic algorithmic improvement as well as other data-type-specific ones over the existing latent variable distillation pipeline. Empirically, we outperform SoTA TPMs by a large margin on challenging image modeling benchmarks. In particular, on ImageNet32, PCs achieve 4.06 bits-per-dimension, which is only 0.34 behind variational diffusion models (Kingma et al., 2021).