HIGH-DIMENSIONAL FACTOR REGRESSION FOR HETEROGENEOUS SUBPOPULATIONS.

HIGH-DIMENSIONAL FACTOR REGRESSION FOR HETEROGENEOUS SUBPOPULATIONS.
复制标题

DOI:
10.5705/ss.202020.0145
复制
发表时间:
2023-01
期刊:
影响因子:
1.4
通讯作者:
Peiyao Wang;Quefeng Li;D. Shen;Yufeng Liu
Peiyao Wang;Quefeng Li;D. Shen;Yufeng Liu
中科院分区:
数学3区
文献类型:
--
作者:
Peiyao Wang;Quefeng Li;D. Shen;Yufeng Liu

文献摘要

相似文献

在现代科学研究中,由于大量的复杂数据,通常观察到数据异质性。我们为具有异质亚群的数据提出了一个因子回归模型。所提出的模型可以表示为异质和均匀术语的分解。异质项是由不同亚群中的潜在因素驱动的。同质项捕获了协变量的共同变化,并在亚群中共享共同的回归系数。我们提出的模型在全球模型与特定组模型之间达到了良好的平衡。全局模型忽略了数据异质性,而小组特异性模型则分别适合每个亚组。我们证明了我们提出的估计器的估计和预测一致性,并表明它的收敛速率比小组特异性和全球模型的收敛率更好。我们表明,估计潜在因素的额外成本在渐近上可以忽略不计,而最小值仍然可以达到。我们通过研究错误指定的小组特异性模型中的预测误差,进一步证明了我们提出的方法的鲁棒性。最后,我们进行仿真研究并分析了来自阿尔茨海默氏病神经影像学计划的数据集和汇总的微阵列数据集,以进一步证明我们提出的因素回归模型的竞争力和解释性。
In modern scientific research, data heterogeneity is commonly observed owing to the abundance of complex data. We propose a factor regression model for data with heterogeneous subpopulations. The proposed model can be represented as a decomposition of heterogeneous and homogeneous terms. The heterogeneous term is driven by latent factors in different subpopulations. The homogeneous term captures common variation in the covariates and shares common regression coefficients across subpopulations. Our proposed model attains a good balance between a global model and a group-specific model. The global model ignores the data heterogeneity, while the group-specific model fits each subgroup separately. We prove the estimation and prediction consistency for our proposed estimators, and show that it has better convergence rates than those of the group-specific and global models. We show that the extra cost of estimating latent factors is asymptotically negligible and the minimax rate is still attainable. We further demonstrate the robustness of our proposed method by studying its prediction error under a mis-specified group-specific model. Finally, we conduct simulation studies and analyze a data set from the Alzheimer's Disease Neuroimaging Initiative and an aggregated microarray data set to further demonstrate the competitiveness and interpretability of our proposed factor regression model.