Bandwidth selection: Classical or plug-in?

Bandwidth selection: Classical or plug-in?
复制标题

DOI:
10.1214/aos/1018031201
复制
发表时间:
1999-04-01
影响因子:
4.5
通讯作者:
Loader, CR
Loader, CR
中科院分区:
数学1区
文献类型:
--
作者:
Loader, CR

文献摘要

被引文献

相似文献

在过去的十年中,用于核密度估计和局部回归等过程的带宽选择得到了广泛的研究。与交叉验证等方法相比,已经收集了大量的证据来建立现代插件方法的优越性能:这包括从详细的收敛速度分析到模拟,再到在真实数据集上的优越性能。在这项工作中,我们详细地查看了其中的一些证据,寻找差异的来源。我们的发现在几个方面挑战了插件方法所宣称的优越性。首先,插件方法严重依赖于任意指定的导频带宽,当该指定错误时,插件方法就会失败。其次,经常被引用的交叉验证的可变性和欠平滑只是反映了带宽选择的不确定性;插件方法通过过度平滑和在给定困难问题时遗漏重要特征来反映这种不确定性。第三,我们来看看渐近理论。插件方法以低效的方式使用可用的曲率信息,从而导致低效的估计。以前与经典方法的比较逐渐地惩罚了经典方法的这种低效,基于插件的估计被它们自己的导频估计所击败。
Bandwidth selection for procedures such as kernel density estimation and local regression have been widely studied over the past decade. Substantial "evidence" has been collected to establish superior performance of modern plug-in methods in comparison to methods such as cross validation: this has ranged from detailed analysis of rates of convergence, to simulations, to superior performance on real datasets.In this work we take a detailed look at some of this evidence, looking into the sources of differences. Our findings challenge the claimed superiority of plug-in methods on several fronts. First, plug-in methods are heavily dependent on arbitrary specification of pilot bandwidths and fail when this specification is wrong. Second, the often-quoted variability and undersmoothing of cross validation simply reflects the uncertainty of bandwidth selection; plug-in methods reflect this uncertainty by oversmoothing and missing important features when given difficult problems. Third, we look at asymptotic theory. Plug-in methods use available curvature information in an inefficient manner, resulting in inefficient estimates. Previous comparisons with classical approaches penalized the classical approaches for this inefficiency Asymptotically, the plug-in based estimates are beaten by their own pilot estimates.