MODEL UNCERTAINTY, DATA MINING AND STATISTICAL-INFERENCE

MODEL UNCERTAINTY, DATA MINING AND STATISTICAL-INFERENCE
复制标题

DOI:
10.2307/2983440
复制
发表时间:
1995-01-01
影响因子:
2
通讯作者:
CHATFIELD, C
CHATFIELD, C
中科院分区:
数学4区
文献类型:
--
作者:
CHATFIELD, C

文献摘要

被引文献

相似文献

本文采用国外实用主义的观点,将统计推理纳入模型制定的各个方面。模型参数的估计传统上假设模型具有预先指定的已知形式,并且不考虑模型结构可能存在的不确定性。这隐含地假设了一个“真实”模式的存在,而许多人会认为这是虚构的。在实践中,模型的不确定性是生活中的事实,可能比其他不确定性来源更为严重,这些不确定性来源受到统计学家的更多关注。无论模型是根据主题指定的,还是以迭代、互动的方式在同一数据集上制定、拟合和检查模型(这种情况越来越多),都是如此。现代计算能力允许考虑大量模型,并且与数据相关的规范搜索已成为许多统计领域的规范。当分析人员竭尽全力获得良好的拟合时,术语数据挖掘可能会在这种上下文中使用。本文回顾了模型不确定性的影响,如过于狭窄的预测区间,以及参数估计中的非平凡偏差,这些偏差可以遵循基于数据的建模。讨论了评估和克服模型不确定性影响的方法,包括使用模拟和重新采样方法,贝叶斯模型平均方法以及尽可能收集额外数据。也许这篇论文的主要目的是确保统计学家意识到这些问题,并开始解决这些问题,即使没有简单的、通用的理论解决方案。
This paper takes abroad, pragmatic view of statistical inference to include all aspects of model formulation. The estimation of model: parameters traditionally assumes that a model has a prespecified known form and takes no account of possible uncertainty regarding the model structure. This implicitly assumes the existence of a 'true' model, which many would regard-as a fiction. In practice model uncertainty is a fact of life and likely to be more serious than other sources of uncertainty which have received far more attention from statisticians. This is true whether the model is specified on subject-matter grounds or, as is increasingly the case, when a model is formulated, fitted and checked on the same data set in an iterative, interactive way. Modern computing power allows a large number of models to be considered and data-dependent specification searches have become the norm in many areas of statistics. The term data mining may be used in this context when the analyst goes to great lengths to obtain a good fit. This paper reviews the effects of model uncertainty, such as too narrow prediction intervals, and the non-trivial biases in parameter estimates which can follow data-based modelling. Ways of assessing and overcoming the effects of model uncertainty are discussed, including the use of simulation and resampling methods, a Bayesian model averaging approach and collecting additional data wherever possible. Perhaps the main aim of the paper is to ensure that statisticians are aware of the problems and start addressing the issues even if there is no simple, general theoretical fix.