Bandwidth selection for kernel density estimators of multivariate level sets and highest density regions

Bandwidth selection for kernel density estimators of multivariate level sets and highest density regions
复制标题

DOI:
10.1214/18-ejs1501
复制
发表时间:
2018-01-01
影响因子:
1.1
通讯作者:
Weng, Guangwei
Weng, Guangwei
中科院分区:
数学3区
文献类型:
--
作者:
Doss, Charles R.;Weng, Guangwei

文献摘要

被引文献

相似文献

本文研究了密度水平集的核密度估计的带宽矩阵选择问题。我们还考虑估计最高密度区域,这不同于估计水平集,因为它指定了集合的概率内容,而不是直接指定水平。这使问题复杂化。KDE的带宽选择已经得到了很好的研究,但大多数方法的目标是最小化密度或其导数的全局损失函数。我们在这里考虑的损失是真实集和估计集的对称差的度量。我们推导出相应风险的渐近近似。近似值取决于可以估计的未知量,然后可以将近似值最小化以产生带宽的选择,我们在模拟中表现良好。我们提供了一个R包lsbs来实现我们的过程。
We consider bandwidth matrix selection for kernel density estimators of density level sets in R-d, d >= 2. We also consider estimation of highest density regions, which differs from estimating level sets in that one specifies the probability content of the set rather than specifying the level directly. This complicates the problem. Bandwidth selection for KDEs is well studied, but the goal of most methods is to minimize a global loss function for the density or its derivatives. The loss we consider here is instead the measure of the symmetric difference of the true set and estimated set. We derive an asymptotic approximation to the corresponding risk. The approximation depends on unknown quantities which can be estimated, and the approximation can then be minimized to yield a choice of bandwidth, which we show in simulations performs well. We provide an R package lsbs for implementing our procedure.