Subjective Bayesian analysis for surveys with missing data
Subjective Bayesian analysis for surveys with missing data
复制标题
对缺失数据的调查进行主观贝叶斯分析
DOI:
--
复制
发表时间:
1993
期刊:
影响因子:
--
通讯作者:
J. Kadane
中科院分区:
文献类型:
--
作者:
J. Kadane
Almost every survey has missing data, sometimes because of inadequacy of the sampling frame, sometimes because of the unwillingness of persons in the frame to participate, sometimes because of the inability of the surveyors to find everyone in the frame, and usually because of an unknown mixture of all these reasons and others. The analyses of surveys often ignore the missing data, treating it as a complete sample of those who responded. However, to do this is often to assume away the principal source of uncertainty, rendering statements of uncertainty conditioned on that assumption problematic. One solution is the assumption of 'ignorability', as discussed by Don Rubin. This approach essentially states in conditional probability terms what one must assume, to get away with an analysis that does not take into account the missing data. The purpose of this paper is to explore a second approach, in which differing beliefs about what the missing data would have been had they been collected are explored to see how robust the results of the survey are. What is reasonable to assume depends on subjective judgements of the analyst. These ideas are considered in the context of a survey on juror death penalty attitudes and behavior. The comparison of subjectivist Bayesian and objectivist frequentistic methods applied to practical problems is very important to the advancement of both methodological positions. This paper is devoted to the consideration of a type of data in which one would have thought that the frequentists would be at their best. The frequency argument is based on viewing a given instance or sample as a member of an infinite stream of independent and identically distributed such instances, and associating the probability of the instance with the relative frequency in the infinite stream. Of course, there are data situations that are very awkward from this stance. Consider, for example, the probability that it will rain in Nottingham sometime tomorrow. Suppose that many years of past data are available. With what subset of this data should I associate tomorrow's rain? Should I choose only those in which the wind, rain and temperature conditions in Dublin match those of today? In making such choices the subjectivity of the associated 'infinite' stream is apparent, and is an embarrassment to a frequentist seeking objectivity. Thus, to explore the possible usefulness of frequentistic ideas, one must choose an example more carefully, in the hope of finding one more congenial to their approach. Surely sampling from a fixed population comes to mind as a candidate. There are two possible senses of sampling that then come to mind: that the individuals in the sample are a random sample from a frame of such individuals, and that the sample itself is random from a frame of subsets, perhaps subsets of the same size. Since the latter results in a sample of size one, it seems very weak in a frequentistic sense. Consequently, the first interpretation is concentrated on here. In a practical implementation of sampling from a human population, not everyone responds. People move, are busy, do not bother and find the questions objectionable. In a reasonably good survey, perhaps 70% of the chosen sample responds. This practical fact imposes a heavy burden on the analysis of the survey data. One response that might be made is to take the sample as representative only of those who would have responded had they been asked. While this has the attraction of objectivistic frequentist ideological purity, it has costs. One is that the sampling frame is of unknown