The number of subjects per variable required in linear regression analyses

The number of subjects per variable required in linear regression analyses
复制标题

DOI:
10.1016/j.jclinepi.2014.12.014
复制
发表时间:
2015-06-01
影响因子:
7.2
通讯作者:
Steyerberg, Ewout W.
Steyerberg, Ewout W.
中科院分区:
医学2区
文献类型:
--
作者:
Austin, Peter C.;Steyerberg, Ewout W.

文献摘要

被引文献

相似文献

目的:为了确定可以包含在线性回归模型中的自变量的数量。研究设计和设置:我们使用一系列Monte Carlo模拟来检查每个变量的受试者数量(SPY)对估计的回归系数和标准误的准确性、估计的置信区间的经验覆盖率以及拟合模型的估计的R-2的准确性的影响。结果:最少约两个SPV往往导致回归系数估计值的相对偏倚小于10%。此外,用这个最小数目的SPY,回归系数的标准误差被准确地估计,并且估计的置信区间近似于广告的覆盖率。尽管调整后的R-2估计值表现良好,但需要更多的SPV来最大限度地减少模型R-2估计值的偏倚。在估计模型R-2统计量的偏差是成反比的人口regression model.Conclusion解释的变异的比例的大小:线性回归模型只需要两个SPV充分估计回归系数,标准误和置信区间。(C)2015作者爱思唯尔公司出版
Objectives: To determine the number of independent variables that can be included in a linear regression model.Study Design and Setting: We used a series of Monte Carlo simulations to examine the impact of the number of subjects per variable (SPY) on the accuracy of estimated regression coefficients and standard errors, on the empirical coverage of estimated confidence intervals, and on the accuracy of the estimated R-2 of the fitted model.Results: A minimum of approximately two SPV tended to result in estimation of regression coefficients with relative bias of less than 10%. Furthermore, with this minimum number of SPY, the standard errors of the regression coefficients were accurately estimated and estimated confidence intervals had approximately the advertised coverage rates. A much higher number of SPV were necessary to minimize bias in estimating the model R-2, although adjusted R-2 estimates behaved well. The bias in estimating the model R-2 statistic was inversely proportional to the magnitude of the proportion of variation explained by the population regression model.Conclusion: Linear regression models require only two SPV for adequate estimation of regression coefficients, standard errors, and confidence intervals. (C) 2015 The Authors. Published by Elsevier Inc.