Estimating Density Models with Truncation Boundaries using Score Matching

Estimating Density Models with Truncation Boundaries using Score Matching
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Song Liu;T. Kanamori;Daniel J. Williams
Song Liu;T. Kanamori;Daniel J. Williams
中科院分区:
其他
文献类型:
--
作者:
Song Liu;T. Kanamori;Daniel J. Williams

文献摘要

相似文献

截断密度是定义在截断域上的概率密度函数。它们与它们的非截断对应物共享相同的参数形式,直到归一化常数。由于其归一化常数的计算通常是不可行的,极大似然估计不容易应用于截断密度模型的估计。分数匹配(SM)是一个功能强大的工具,只使用非标准化模型拟合参数。然而,它不能直接应用于这里作为边界条件,用于导出一个听话的SM目标不满足截断密度。本文研究了截断概率密度的参数估计问题。估计量最小化加权Fisher散度。权函数是从数据点到域边界的最短距离。我们表明,这种选择的权重函数自然产生于最小化Stein差异以及上界有限样本估计误差。数值实验和对芝加哥犯罪数据集的研究证明了该方法的有效性。我们还表明,建议的密度估计可以纠正异常值检测方法所造成的异常值修剪偏差。
Truncated densities are probability density functions defined on truncated domains. They share the same parametric form with their non-truncated counterparts up to a normalizing constant. Since the computation of their normalizing constants is usually infeasible, Maximum Likelihood Estimation cannot be easily applied to estimate truncated density models. Score Matching (SM) is a powerful tool for fitting parameters using only unnormalized models. However, it cannot be directly applied here as boundary conditions used to derive a tractable SM objective are not satisfied by truncated densities. In this paper, we study parameter estimation for truncated probability densities using SM. The estimator minimizes a weighted Fisher divergence. The weight function is simply the shortest distance from a data point to the boundary of the domain. We show this choice of weight function naturally arises from minimizing the Stein discrepancy as well as upperbounding the finite-sample estimation error. The usefulness of our method is demonstrated by numerical experiments and a study on the Chicago crime data set. We also show that the proposed density estimation can correct the outlier-trimming bias caused by aggressive outlier detection methods.