Robust Kullback-Leibler Divergence and Universal Hypothesis Testing for Continuous Distributions

Robust Kullback-Leibler Divergence and Universal Hypothesis Testing for Continuous Distributions
复制标题

DOI:
10.1109/tit.2018.2879057
复制
发表时间:
2017-11
影响因子:
2.5
通讯作者:
Pengfei Yang;Biao Chen
Pengfei Yang;Biao Chen
中科院分区:
计算机科学2区
文献类型:
--
作者:
Pengfei Yang;Biao Chen

文献摘要

相似文献

通用假设检验(UHT)是指决定样本是否来自标称分布或与标称分布不同的未知分布的问题。Hoeffding检验,其检验统计量相当于经验Kullback-Leibler发散(KL发散),已知对于定义在有限字母表上的分布是渐近最优的。然而,随着连续观测,分布函数中KL发散的不连续性导致UHT的显着并发症。本文介绍了经典KL散度的一个鲁棒版本,定义为从一个分布到一个已知分布的Lévy球的KL散度。这种强大的KL分歧被证明是连续的基本分布函数的弱收敛。连续性的性质,使大学假设检验问题的连续观测的渐近最优检验的发展。其最优性与Hoeffding检验的最优性意义相同,但比Zeitouni和Gutman检验的最优性意义更强。也许更重要的是,开发的测试统计量可以通过凸程序计算,使其在实践中更有意义。数值实验也进行了评估其性能相比,最近提出的一些基于核的拟合优度测试。
Universal hypothesis testing (UHT) refers to the problem of deciding whether samples come from a nominal distribution or an unknown distribution that is different from the nominal distribution. Hoeffding’s test, whose test statistic is equivalent to the empirical Kullback–Leibler divergence (KL divergence), is known to be asymptotically optimal for distributions defined on finite alphabets. With continuous observations, however, the discontinuity of the KL divergence in the distribution functions results in significant complications for UHT. This paper introduces a robust version of the classical KL divergence, defined as the KL divergence from a distribution to the Lévy ball of a known distribution. This robust KL divergence is shown to be continuous in the underlying distribution function with respect to the weak convergence. The continuity property enables the development of an asymptotically optimal test for the university hypothesis testing problem with continuous observations. The optimality is in the same sense as that of the Hoeffding’s test and stronger than that of Zeitouni and Gutman. Perhaps more importantly, the developed test statistic can be computed through convex programs, making it much more meaningful in practice. Numerical experiments are also conducted to evaluate its performance as compared with some kernel based goodness of fit test that has been proposed recently.