POST-STRATIFICATION - A MODELERS PERSPECTIVE

POST-STRATIFICATION - A MODELERS PERSPECTIVE
复制标题

DOI:
10.2307/2290792
复制
发表时间:
1993-09-01
影响因子:
3.7
通讯作者:
LITTLE, RJA
LITTLE, RJA
中科院分区:
数学1区
文献类型:
--
作者:
LITTLE, RJA

文献摘要

被引文献

相似文献

后分层是调查分析中的一种常用技术,用于将变量的人口分布纳入调查估计数。基本技术将样本划分为后层,并为后层h中的每个样本案例计算后层权重w(h)= rP(h)/r(h),其中r(h)是后层h中的调查受访者人数,P(h)是人口普查的人口比例,r是受访者样本量。调查估计,例如均值和总数的函数,然后用w(h)加权。该方法的变体和扩展包括截断权重以避免过度的可变性和对两个或更多个单变量边缘分布的集合的倾斜。关于后分层的文献有限,并且主要采用随机化(或基于设计)的观点,其中推断基于人群值保持固定的抽样分布。本文发展了基于贝叶斯模型的理论方法。一个基本的正常后分层模型,产生后分层的平均值作为后均值,和后方差,包括估计方差的调整。然后提出了小样本推断的修改,基于(a)改变后地层参数的Jeffreys先验以跨后地层借用强度,和(B)忽略关于后地层的部分信息。特别是,实际规则崩溃后地层,以减少后方差的开发和频率论的方法相比。两个后分层变量的方法也被认为是。耙样本计数和受访者计数提供近似贝叶斯推断时,利润率的两个后分层,但他们的联合分布是不。当联合分布可用时,raking有效地忽略了它包含的信息,因此可以与其他忽略信息的技术(如折叠)进行比较。对于推断的手段,它建议,耙是最合适的时,后地层的手段有一个添加剂或近添加剂的结构,而崩溃时,相互作用是存在的。
Post-stratification is a common technique in survey analysis for incorporating population distributions of variables into survey estimates. The basic technique divides the sample into post-strata, and computes a post-stratification weight w(h) = rP(h)/r(h) for each sample case in post-stratum h, where r(h) is the number of survey respondents in post-stratum h, P(h) is the population proportion from a census, and r is the respondent sample size. Survey estimates, such as functions of means and totals, then weight cases by w(h). Variants and extensions of the method include truncation of the weights to avoid excessive variability and raking to a set of two or more univariate marginal distributions. Literature on post-stratification is limited and has mainly taken the randomization (or design-based) perspective, where inference is based on the sampling distribution with population values held fixed. This article develops Bayesian model-based theory for the method. A basic normal post-stratification model is introduced which yields the post-stratified mean as the posterior mean, and a posterior variance that incorporates adjustments for estimating variances. Modifications are then proposed for small sample inference, based on (a) changing the Jeffreys prior for the post-stratum parameters to borrow strength across post-strata, and (b) ignoring partial information about the post-strata. In particular, practical rules for collapsing post-strata to reduce posterior variance are developed and compared with frequentist approaches. Methods for two post-stratifying variables are also considered. Raking sample counts and respondent counts is shown to provide approximate Bayesian inferences when the margins of the two post-stratifiers are available but their joint distribution is not. When the joint distribution is available, raking effectively ignores the information it contains, and hence can be compared with other techniques that ignore information such as collapsing. For inference about means, it is suggested that raking is most appropriate when post-stratum means have an additive or near-additive structure, whereas collapsing is indicated when interactions are present.