Editorial: Model selection and efficiency—is ‘Which model …?’ the right question?
Editorial: Model selection and efficiency—is ‘Which model …?’ the right question?
复制标题
DOI:
10.1111/j.1467-985x.2005.00366.x
复制
发表时间:
2005-07
期刊:
影响因子:
--
通讯作者:
N. Longford
中科院分区:
文献类型:
--
作者:
N. Longford
Statistics as a science and profession has been transformed over the last few decades by finetuning its orientation to serve other scientific fields, encouraged by the revolution in computing technology. As a result, it encompasses a vast variety of activities, provides a wide range of careers and is universally accepted as indispensable to the modern information society. These positive aspects go hand in hand with profound weaknesses. We cannot agree on an authoritative definition of our subject, on a short list of its fundamental principles (those of probability theory are insufficient) or on what amounts to good practice in particular settings, and how to promote it. Model selection, with the associated uncertainty, is an example of practice that is becoming increasingly problematic as powerful computers and convenient software enable us to explore data in ever greater detail. We can inspect how several alternative models fit the studied data set, and settle on one of them. Such a ‘final’ model and its maximum likelihood (ML) fit (estimates and standard errors, or information equivalent to them) is the centre-piece of the results section of many a report or manuscript, accompanied by model checking and claims of (approximate) unbiasedness and asymptotic efficiency, after confirming that the requisite regularity conditions have been satisfied. Despite being regarded as respectable, this approach is flawed because it ignores the consequences of model uncertainty. Since the lucid discussions by Draper (1995) and Chatfield (1995), neither research nor practice has paid much attention to this issue. To Bayesians, the topic might be broached constructively by paraphrasing de Finnetti (1974) (‘Every probability is conditional’) as ‘Every posterior distribution is conditional’. We usually study the properties of estimators conditionally on the selected model, ruling out the possibility that the selected model might not be valid. After all, the model selection is a random (data-dependent) process, but the (unknown) ‘good’ model is fixed, being a property of the studied phenomenon and oblivious to our study design and data collection process. Model selection forces us to make a decision without considering the consequences of the errors that may have been made in the process. We end up putting all our inferential eggs in one unevenly woven basket. Depending on the purpose, an error of one kind may be innocuous or disastrous relative to an error of another kind. The conditional probabilities of these two kinds of error, controlled in hypothesis testing, are often a poor indication of their gravity. By way of an example, consider the text-book balanced one-way analysis-of-variance (ANOVA) setting (normality and equal variances within groups) with K = 8 observations in each of J =10 groups, and two problems: