A note on centering in subsample selection for linear regression

A note on centering in subsample selection for linear regression
复制标题

DOI:
10.1002/sta4.525
复制
发表时间:
2022-09
期刊:
影响因子:
1.7
通讯作者:
Hai Ying Wang
Hai Ying Wang
中科院分区:
数学4区
文献类型:
--
作者:
Hai Ying Wang

文献摘要

被引文献

相似文献

居中是线性回归分析中常用的技术。使用响应和协变量的中心数据,可以从没有截距的模型计算斜率参数的普通最小二乘估计量。如果从居中的完整数据中选择子样本,则该子样本通常是非居中的。在这种情况下,在没有截距的情况下拟合模型仍然合适吗?答案是肯定的,我们表明,最小二乘估计的斜率参数从一个模型没有截距是无偏的,它有一个较小的方差协方差矩阵的Loewner顺序比从一个模型的截距。我们进一步表明,对于非信息加权子采样时,使用加权最小二乘估计,使用全数据加权均值重新定位子样本提高了估计效率。
Centring is a commonly used technique in linear regression analysis. With centred data on both the responses and covariates, the ordinary least squares estimator of the slope parameter can be calculated from a model without the intercept. If a subsample is selected from a centred full data, the subsample is typically uncentred. In this case, is it still appropriate to fit a model without the intercept? The answer is yes, and we show that the least squares estimator on the slope parameter obtained from a model without the intercept is unbiased and it has a smaller variance covariance matrix in the Loewner order than that obtained from a model with the intercept. We further show that for noninformative weighted subsampling when a weighted least squares estimator is used, using the full data weighted means to relocate the subsample improves the estimation efficiency.