Statistical Efficiency of Score Matching: The View from Isoperimetry

Statistical Efficiency of Score Matching: The View from Isoperimetry
复制标题

DOI:
10.48550/arxiv.2210.00726
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Frederic Koehler;Alexander Heckett;Andrej Risteski
Frederic Koehler;Alexander Heckett;Andrej Risteski
中科院分区:
其他
文献类型:
--
作者:
Frederic Koehler;Alexander Heckett;Andrej Risteski

文献摘要

相似文献

参数化到归一化常数的深度生成模型(例如,基于能量的模型)难以通过最大化数据的似然来训练,因为其似然和/或梯度不能被显式地或有效地写下来。分数匹配是一种训练方法,它不是为训练数据拟合似然$\log p(x)$,而是拟合分数函数$\nabla_x \log p(x)$ -避免了评估分区函数的需要。虽然这个估计量已知是一致的,但它的统计效率是否(以及何时)与最大似然法相当还不清楚,最大似然法已知是(渐近)最优的。我们在本文中启动了这条调查线,并显示了统计效率的得分匹配和等周属性的分布估计-即庞加莱,对数Sobolev和等周常数-数量管理的混合时间的马尔可夫过程,如朗之万动力学之间的紧密联系。粗略地说,我们表明,分数匹配估计是统计上可比的最大似然分布时,有一个小的等周常数。相反,如果分布有一个大的等周常数-即使是简单的分布族,如指数族,具有足够丰富的统计数据-分数匹配将大大低于最大似然法的效率。我们适当地形式化这些结果在有限样本制度,并在渐近制度。最后,我们确定了一个直接平行的离散设置,在那里我们连接的统计特性的pseudolikkastrointestimation近似张量熵和Glauber动力学。
Deep generative models parametrized up to a normalizing constant (e.g. energy-based models) are difficult to train by maximizing the likelihood of the data because the likelihood and/or gradients thereof cannot be explicitly or efficiently written down. Score matching is a training method, whereby instead of fitting the likelihood $\log p(x)$ for the training data, we instead fit the score function $\nabla_x \log p(x)$ -- obviating the need to evaluate the partition function. Though this estimator is known to be consistent, its unclear whether (and when) its statistical efficiency is comparable to that of maximum likelihood -- which is known to be (asymptotically) optimal. We initiate this line of inquiry in this paper, and show a tight connection between statistical efficiency of score matching and the isoperimetric properties of the distribution being estimated -- i.e. the Poincar\'e, log-Sobolev and isoperimetric constant -- quantities which govern the mixing time of Markov processes like Langevin dynamics. Roughly, we show that the score matching estimator is statistically comparable to the maximum likelihood when the distribution has a small isoperimetric constant. Conversely, if the distribution has a large isoperimetric constant -- even for simple families of distributions like exponential families with rich enough sufficient statistics -- score matching will be substantially less efficient than maximum likelihood. We suitably formalize these results both in the finite sample regime, and in the asymptotic regime. Finally, we identify a direct parallel in the discrete setting, where we connect the statistical properties of pseudolikelihood estimation with approximate tensorization of entropy and the Glauber dynamics.