A Witness Function Based Construction of Discriminative Models Using Hermite Polynomials

A Witness Function Based Construction of Discriminative Models Using Hermite Polynomials
复制标题

DOI:
10.3389/fams.2020.00031
复制
发表时间:
2019-01
期刊:
--
影响因子:
--
通讯作者:
H. Mhaskar;A. Cloninger;Xiuyuan Cheng
H. Mhaskar;A. Cloninger;Xiuyuan Cheng
中科院分区:
其他
文献类型:
--
作者:
H. Mhaskar;A. Cloninger;Xiuyuan Cheng

文献摘要

被引文献

相似文献

在机器学习中,我们获得了{(xj,yj)} j = 1m的数据集,绘制为i.i.d。类μk*(x)可能重叠而不是使用正核,例如高斯这些度量的估计,使用保留大量这些度量矩的非阳性内核产生最佳的近似值。概率的意义与相同的内核开发的置换测试,我们证明了内核估计器在分类问题中成为“证人功能”。该点x的估计器超过一定阈值,那么该方法在某个类别中可靠地可靠。 - 出于通用模型,分类不确定性或查找强大的质心的目的。新闻文件和拉隆德数据集。
In machine learning, we are given a dataset of the form {(xj,yj)}j=1M, drawn as i.i.d. samples from an unknown probability distribution μ; the marginal distribution for the xj's being μ*, and the marginals of the kth class μk*(x) possibly overlapping. We address the problem of detecting, with a high degree of certainty, for which x we have μk*(x)>μi*(x) for all i ≠ k. We propose that rather than using a positive kernel such as the Gaussian for estimation of these measures, using a non-positive kernel that preserves a large number of moments of these measures yields an optimal approximation. We use multi-variate Hermite polynomials for this purpose, and prove optimal and local approximation results in a supremum norm in a probabilistic sense. Together with a permutation test developed with the same kernel, we prove that the kernel estimator serves as a “witness function” in classification problems. Thus, if the value of this estimator at a point x exceeds a certain threshold, then the point is reliably in a certain class. This approach can be used to modify pretrained algorithms, such as neural networks or nonlinear dimension reduction techniques, to identify in-class vs out-of-class regions for the purposes of generative models, classification uncertainty, or finding robust centroids. This fact is demonstrated in a number of real world data sets including MNIST, CIFAR10, Science News documents, and LaLonde data sets.