What you see may not be what you get: A brief, nontechnical introduction to overfitting in regression-type models

What you see may not be what you get: A brief, nontechnical introduction to overfitting in regression-type models
复制标题

DOI:
10.1097/01.psy.0000127692.23278.a9
复制
发表时间:
2004-05-01
影响因子:
3.3
通讯作者:
Babyak, MA
Babyak, MA
中科院分区:
医学3区
文献类型:
--
作者:
Babyak, MA

文献摘要

被引文献

相似文献

目的:统计模型,如线性或逻辑回归或生存分析,经常被用来回答心身研究中的科学问题。然而,许多使用这些技术的人显然没有充分认识到过度拟合的问题,即利用手头样本的特性。过度拟合的模型将无法在未来的样本中复制,从而对该发现的科学价值产生相当大的不确定性。本文是对过拟合概念的非技术性讨论,旨在让具有不同统计专业知识水平的读者能够理解。过拟合的概念是从现有数据中要求太多。给定一个数据集中一定数量的观测值,模型的复杂性有一个上限,可以用任何可接受的不确定性程度来推导。复杂性是数据分析任何阶段针对同一数据集所扩展的自由度(包括相互作用和非线性项等复杂项的预测因子数量)的函数。理论和经验证据,特别注重计算机模拟研究的结果,以证明过拟合的实际后果与科学推理。三种常见的做法自动变量选择,候选预测因子的预测试,和二分法的连续变量,是造成相当大的风险,在模型中的虚假结果。过拟合和探索候选混杂因素之间的困境也进行了讨论。讨论了防止过拟合的其他方法,包括变量聚合和先验系数的固定。还介绍了计算和纠正复杂性的技术,包括收缩和惩罚。
Objective: Statistical models, such as linear or logistic regression or survival analysis, are frequently used as a means to answer scientific questions in psychosomatic research. Many who use these techniques, however, apparently fail to appreciate fully the problem of overfitting, ie, capitalizing on the idiosyncrasies of the sample at hand. Overfitted models will fail to replicate in future samples, thus creating considerable uncertainty about the scientific merit of the finding. The present article is a nontechnical discussion of the concept of overfitting and is intended to be accessible to readers with varying levels of statistical expertise. The notion of overfitting is presented in terms of asking too much from the available data. Given a certain number of observations in a data set, there is an upper limit to the complexity of the model that can be derived with any acceptable degree of uncertainty. Complexity arises as a function of the number of degrees of freedom expended (the number of predictors including complex terms such as interactions and nonlinear terms) against the same data set during any stage of the data analysis. Theoretical and empirical evidence-with a special focus on the results of computer simulation studies-is presented to demonstrate the practical consequences of overfitting with respect to scientific inference. Three common practices-automated variable selection, pretesting of candidate predictors, and dichotomization of continuous variables-are shown to pose a considerable risk for spurious findings in models. The dilemma between overfitting and exploring candidate confounders is also discussed. Alternative means of guarding against overfitting are discussed, including variable aggregation and the fixing of coefficients a priori. Techniques that account and correct for complexity, including shrinkage and penalization, also are introduced.