Classification Logit Two-Sample Testing by Neural Networks for Differentiating Near Manifold Densities

Classification Logit Two-Sample Testing by Neural Networks for Differentiating Near Manifold Densities
复制标题

DOI:
10.1109/tit.2022.3175691
复制
发表时间:
2019-09
影响因子:
2.5
通讯作者:
Xiuyuan Cheng;A. Cloninger
Xiuyuan Cheng;A. Cloninger
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xiuyuan Cheng;A. Cloninger

文献摘要

相似文献

生成对抗网络和变分学习的最新成功表明,培训分类网络可能在解决经典的两样本问题方面可以很好地奏效,该问题要求鉴于每个样本的有限样本来区分两种密度。基于网络的方法具有计算优势,即该算法会扩展到大型数据集。本文考虑使用分类logit函数,该函数由训练有素的分类神经网络提供,并在两个数据集的测试集拆分上进行了评估,以计算两个样本统计量。为了分析logit函数的近似和估计误差以区分近型密度,我们引入了神经网络近曼佛积分近似的新结果。然后,我们证明logit函数可证明可以区分两个亚指数密度,因为该网络已充分参数化,并且对于ON或接近歧管密度,所需的网络复杂性仅随着固有维度而缩小为缩小。在实验中,网络Logit测试使用分类精度表现出比以前基于网络的测试更好的性能,并且在合成数据集和手写数字数据集上的某些内核最大均值差异测试中也有利。
The recent success of generative adversarial networks and variational learning suggests that training a classification network may work well in addressing the classical two-sample problem, which asks to differentiate two densities given finite samples from each one. Network-based methods have the computational advantage that the algorithm scales to large datasets. This paper considers using the classification logit function, which is provided by a trained classification neural network and evaluated on the testing set split of the two datasets, to compute a two-sample statistic. To analyze the approximation and estimation error of the logit function to differentiate near-manifold densities, we introduce a new result of near-manifold integral approximation by neural networks. We then show that the logit function provably differentiates two sub-exponential densities given that the network is sufficiently parametrized, and for on or near manifold densities, the needed network complexity is reduced to only scale with the intrinsic dimensionality. In experiments, the network logit test demonstrates better performance than previous network-based tests using classification accuracy, and also compares favorably to certain kernel maximum mean discrepancy tests on synthetic datasets and hand-written digit datasets.