A valid and fast spatial bootstrap for correlation functions

A valid and fast spatial bootstrap for correlation functions
复制标题

DOI:
10.1086/588631
复制
发表时间:
2008-07-01
影响因子:
4.9
通讯作者:
Loh, Ji Meng
Loh, Ji Meng
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Loh, Ji Meng

文献摘要

被引文献

相似文献

在本文中,我们研究的有效性非参数空间自助作为一个程序来量化的N点相关函数的估计误差。我们这样做是通过一个小的模拟研究与简单的点过程模型和估计的两点相关函数及其误差。使用自举得到的置信区间的覆盖范围进行了比较,从假设泊松误差。这里考虑的自举过程适于与空间(即,依赖)数据。特别是,我们描述了一个标记点的引导程序,而不是重新采样点或点块,我们重新采样分配给数据点的标记。这些标记是基于感兴趣的统计数据的数值。我们描述了如何为两点和三点相关函数定义标记。通过重新标记,引导样本保留了数据中存在的更多依赖结构。此外,这种引导方法可以比其他一些空间数据的引导方法更快地执行,使其成为一种更实用的方法与大数据集。我们发现,与聚类点数据集,使用标记点自举得到的置信区间的经验覆盖率更接近标称水平比使用泊松误差得到的置信区间。自举误差也被发现更接近于聚类点数据集的真实误差。
In this paper we examine the validity of nonparametric spatial bootstrap as a procedure to quantify errors in estimates of N-point correlation functions. We do this by means of a small simulation study with simple point process models and estimating the two-point correlation functions and their errors. The coverage of confidence intervals obtained using bootstrap is compared with those obtained from assuming Poisson errors. The bootstrap procedure considered here is adapted for use with spatial (i.e., dependent) data. In particular, we describe a marked point bootstrap where, instead of resampling points or blocks of points, we resample marks assigned to the data points. These marks are numerical values that are based on the statistic of interest. We describe how the marks are defined for the two- and three-point correlation functions. By resampling marks, the bootstrap samples retain more of the dependence structure present in the data. Furthermore, this method of bootstrap can be performed much quicker than some other bootstrap methods for spatial data, making it a more practical method with large data sets. We find that with clustered point data sets, confidence intervals obtained using the marked point bootstrap has empirical coverage closer to the nominal level than the confidence intervals obtained using Poisson errors. The bootstrap errors were also found to be closer to the true errors for the clustered point data sets.