Descriptive inference using large, unrepresentative nonprobability samples: An introduction for ecologists.

Descriptive inference using large, unrepresentative nonprobability samples: An introduction for ecologists.
复制标题

使用大型、不具代表性的非概率样本进行描述性推理:生态学家简介。

DOI:
10.1002/ecy.4214
复制
发表时间:
2024
期刊:
影响因子:
4.8
通讯作者:
Boyd RJ
Boyd RJ
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Boyd RJ

文献摘要

相似文献

生物多样性监测通常涉及根据在该景观内取样地点所作的观察,对特定景观中某些感兴趣的变量作出推断。如果感兴趣的变量在采样和非采样位置之间不同,并且没有采取缓解措施,那么样本是不具代表性的,并且从中得出的推论将是有偏差的。有可能调整不具代表性的样本,使它们更接近于“辅助变量”方面的更广泛的景观。一个好的辅助变量是样本包含和感兴趣变量的常见原因,如果它解释了两者中相当一部分的方差,那么从调整后的样本中得出的推论将更接近事实。我们应用了六种类型的调查样本调整——亚抽样、准随机化、后分层、超种群模型、“双鲁棒”过程、多水平回归和后分层——来解决一个简单的两部分生物多样性监测问题。第一部分估算了1987-1999年和2010-2019年两个时期英国马蹄莲植物的平均占用率;第二步是估计两者之间的差异(即趋势)。我们使用来自公民科学数据集的大量但(最初)不具代表性的样本来估计平均值和趋势。与未经调整的估计值相比,使用大多数调整方法估计的平均值和趋势更准确,尽管标准不确定区间通常不包括真实值。如果不知道并拥有所有相关辅助变量的数据,就不可能从一个不具代表性的样本中做出完全无偏的推断。如果辅助变量可用并仔细选择,调整可以减少偏倚,但应承认和报告潜在的残余偏倚。
Biodiversity monitoring usually involves drawing inferences about some variable of interest across a defined landscape from observations made at a sample of locations within that landscape. If the variable of interest differs between sampled and nonsampled locations, and no mitigating action is taken, then the sample is unrepresentative and inferences drawn from it will be biased. It is possible to adjust unrepresentative samples so that they more closely resemble the wider landscape in terms of “auxiliary variables.” A good auxiliary variable is a common cause of sample inclusion and the variable of interest, and if it explains an appreciable portion of the variance in both, then inferences drawn from the adjusted sample will be closer to the truth. We applied six types of survey sample adjustment—subsampling, quasirandomization, poststratification, superpopulation modeling, a “doubly robust” procedure, and multilevel regression and poststratification—to a simple two‐part biodiversity monitoring problem. The first part was to estimate the mean occupancy of the plantCalluna vulgarisin Great Britain in two time periods (1987–1999 and 2010–2019); the second was to estimate the difference between the two (i.e., the trend). We estimated the means and trend using large, but (originally) unrepresentative, samples from a citizen science dataset. Compared with the unadjusted estimates, the means and trends estimated using most adjustment methods were more accurate, although standard uncertainty intervals generally did not cover the true values. Completely unbiased inference is not possible from an unrepresentative sample without knowing and having data on all relevant auxiliary variables. Adjustments can reduce the bias if auxiliary variables are available and selected carefully, but the potential for residual bias should be acknowledged and reported.