Optimal Nonparametric Inference with Two-Scale Distributional Nearest Neighbors

Optimal Nonparametric Inference with Two-Scale Distributional Nearest Neighbors
复制标题

具有两尺度分布最近邻的最优非参数推理

DOI:
10.1080/01621459.2022.2115375
复制
发表时间:
2022
影响因子:
3.7
通讯作者:
Wang, Jingbo
Wang, Jingbo
中科院分区:
数学1区
文献类型:
--
作者:
Demirkaya, Emre;Fan, Yingying;Gao, Lan;Lv, Jinchi;Vossler, Patrick;Wang, Jingbo

文献摘要

相似文献

加权最近邻(WNN)估计作为一种灵活且易于实现的非参数均值回归估计工具已被广泛应用。装袋技术是一种优雅的方法来形成WNN估计量,其权重自动生成到最近的邻居(Steele ; Biau,Cérou和Guyader);我们将所得估计量命名为分布最近邻居(DNN),以便于参考。然而,这种估计量缺乏分布结果,限制了它在统计推断中的应用。此外,当均值回归函数具有高阶光滑性时,DNN不会达到最佳的非参数收敛速度,主要是因为偏差问题。在这项工作中,我们对DNN进行了深入的技术分析,在此基础上,我们通过线性组合两个具有不同子采样尺度的DNN估计器,提出了DNN估计器的偏倚减少方法,从而产生了新的双尺度DNN(TDNN)估计器。双尺度DNN估计器具有WNN的等价表示,其权重允许显式形式,并且一些权重为负。我们证明了,由于使用负权重,双尺度DNN估计在四阶光滑条件下估计回归函数时具有最佳非参数收敛速度。我们进一步超越了估计,并确定DNN和双尺度DNN都是渐进正态的,因为子采样尺度和样本量偏离到无穷大。对于实际的实现,我们还提供了方差估计和分布估计使用刀切和自举技术的双尺度DNN。这些估计量可以用来构造回归函数的非参数推断的有效置信区间。理论结果和吸引人的有限样本性能的建议的双尺度DNN方法的几个仿真例子和真实的数据应用说明。
The weighted nearest neighbors (WNN) estimator has been popularly used as a flexible and easy-to-implement nonparametric tool for mean regression estimation. The bagging technique is an elegant way to form WNN estimators with weights automatically generated to the nearest neighbors (Steele ; Biau, Cérou, and Guyader ); we name the resulting estimator as the distributional nearest neighbors (DNN) for easy reference. Yet, there is a lack of distributional results for such estimator, limiting its application to statistical inference. Moreover, when the mean regression function has higher-order smoothness, DNN does not achieve the optimal nonparametric convergence rate, mainly because of the bias issue. In this work, we provide an in-depth technical analysis of the DNN, based on which we suggest a bias reduction approach for the DNN estimator by linearly combining two DNN estimators with different subsampling scales, resulting in the novel two-scale DNN (TDNN) estimator. The two-scale DNN estimator has an equivalent representation of WNN with weights admitting explicit forms and some being negative. We prove that, thanks to the use of negative weights, the two-scale DNN estimator enjoys the optimal nonparametric rate of convergence in estimating the regression function under the fourth-order smoothness condition. We further go beyond estimation and establish that the DNN and two-scale DNN are both asymptotically normal as the subsampling scales and sample size diverge to infinity. For the practical implementation, we also provide variance estimators and a distribution estimator using the jackknife and bootstrap techniques for the two-scale DNN. These estimators can be exploited for constructing valid confidence intervals for nonparametric inference of the regression function. The theoretical results and appealing finite-sample performance of the suggested two-scale DNN method are illustrated with several simulation examples and a real data application.