Subjective Bayesian analysis for surveys with missing data

Subjective Bayesian analysis for surveys with missing data
复制标题

对缺失数据的调查进行主观贝叶斯分析

DOI:
--
复制
发表时间:
1993
期刊:
影响因子:
--
通讯作者:
J. Kadane
J. Kadane
中科院分区:
--
文献类型:
--
作者:
J. Kadane

文献摘要

被引文献

相似文献

几乎每一次调查都有缺失的数据,有时是因为抽样框架不足,有时是因为框架中的人不愿意参与,有时是因为调查员无法找到框架中的每个人,通常是因为所有这些原因和其他原因的未知混合。调查分析往往忽略缺失的数据,将其视为回复者的完整样本。然而,这样做往往是假设不确定的主要来源,使以这一假设为条件的不确定陈述成为问题。正如唐·鲁宾所讨论的,一种解决方案是“可知性”的假设。这种方法实质上是用条件概率术语陈述了一个人必须假设的东西,以便在没有考虑缺失数据的分析中脱身。本文的目的是探索第二种方法,即探索不同的信念,即如果收集丢失的数据,它们将是什么,以查看调查结果有多稳健。什么是合理的假设取决于分析师的主观判断。这些想法是在一项关于陪审员死刑态度和行为的调查的背景下考虑的。比较主观主义的贝叶斯方法和客观主义的频度方法在实际问题中的应用,对于这两种方法论立场的提升都是非常重要的。这篇文章致力于考虑一种数据类型,在这种数据中,人们会认为频率会处于最佳状态。频率参数基于将给定实例或样本视为独立且相同分布的这种实例的无限大流的成员,并将该实例的概率与无限大流中的相对频率相关联。当然,从这种立场来看,也有一些数据情况非常尴尬。例如,考虑诺丁汉明天某个时候下雨的可能性。假设过去很多年的数据都是可用的。我应该将这些数据的哪个子集与明天的降雨相关联?我应该只选择都柏林的风、雨和温度条件与今天相匹配的那些吗?在做出这样的选择时,相关的“无限”流的主观性是显而易见的,对于寻求客观性的常客来说,这是一种尴尬。因此,为了探索频繁思想的可能用处,人们必须更仔细地选择一个例子,希望找到一个更适合他们的方法的例子。当然,作为候选人,人们会想到从固定人群中抽样。然后想到两种可能的抽样意义:样本中的个体是来自这样的个体的帧的随机样本,以及样本本身是来自子集的帧的随机样本,可能是相同大小的子集。由于后者的结果是一个大小为1的样本,它在频率意义上似乎非常弱。因此,第一种解释集中在这里。在一个从人群中抽样的实际实施中,并不是每个人都会做出反应。人们搬家,忙碌,不用费心,觉得这些问题令人反感。在一项相当不错的调查中,可能有70%的被选样本做出了回应。这一现实给调查数据的分析带来了沉重的负担。一种可能的回答是,只将样本作为那些如果被问到就会回答的人的代表。虽然这具有客观主义的频繁主义意识形态纯洁性的吸引力,但它也有代价。其一是采样帧未知
Almost every survey has missing data, sometimes because of inadequacy of the sampling frame, sometimes because of the unwillingness of persons in the frame to participate, sometimes because of the inability of the surveyors to find everyone in the frame, and usually because of an unknown mixture of all these reasons and others. The analyses of surveys often ignore the missing data, treating it as a complete sample of those who responded. However, to do this is often to assume away the principal source of uncertainty, rendering statements of uncertainty conditioned on that assumption problematic. One solution is the assumption of 'ignorability', as discussed by Don Rubin. This approach essentially states in conditional probability terms what one must assume, to get away with an analysis that does not take into account the missing data. The purpose of this paper is to explore a second approach, in which differing beliefs about what the missing data would have been had they been collected are explored to see how robust the results of the survey are. What is reasonable to assume depends on subjective judgements of the analyst. These ideas are considered in the context of a survey on juror death penalty attitudes and behavior. The comparison of subjectivist Bayesian and objectivist frequentistic methods applied to practical problems is very important to the advancement of both methodological positions. This paper is devoted to the consideration of a type of data in which one would have thought that the frequentists would be at their best. The frequency argument is based on viewing a given instance or sample as a member of an infinite stream of independent and identically distributed such instances, and associating the probability of the instance with the relative frequency in the infinite stream. Of course, there are data situations that are very awkward from this stance. Consider, for example, the probability that it will rain in Nottingham sometime tomorrow. Suppose that many years of past data are available. With what subset of this data should I associate tomorrow's rain? Should I choose only those in which the wind, rain and temperature conditions in Dublin match those of today? In making such choices the subjectivity of the associated 'infinite' stream is apparent, and is an embarrassment to a frequentist seeking objectivity. Thus, to explore the possible usefulness of frequentistic ideas, one must choose an example more carefully, in the hope of finding one more congenial to their approach. Surely sampling from a fixed population comes to mind as a candidate. There are two possible senses of sampling that then come to mind: that the individuals in the sample are a random sample from a frame of such individuals, and that the sample itself is random from a frame of subsets, perhaps subsets of the same size. Since the latter results in a sample of size one, it seems very weak in a frequentistic sense. Consequently, the first interpretation is concentrated on here. In a practical implementation of sampling from a human population, not everyone responds. People move, are busy, do not bother and find the questions objectionable. In a reasonably good survey, perhaps 70% of the chosen sample responds. This practical fact imposes a heavy burden on the analysis of the survey data. One response that might be made is to take the sample as representative only of those who would have responded had they been asked. While this has the attraction of objectivistic frequentist ideological purity, it has costs. One is that the sampling frame is of unknown