Estimation and Inference for High-Dimensional Generalized Linear Models with Knowledge Transfer

Estimation and Inference for High-Dimensional Generalized Linear Models with Knowledge Transfer
复制标题

具有知识转移的高维广义线性模型的估计与推断

DOI:
10.1080/01621459.2023.2184373
复制
发表时间:
2023-02
影响因子:
3.7
通讯作者:
Sai Li;Linjun Zhang;T. Cai;Hongzhe Li
Sai Li;Linjun Zhang;T. Cai;Hongzhe Li
中科院分区:
数学1区
文献类型:
--
作者:
Sai Li;Linjun Zhang;T. Cai;Hongzhe Li

文献摘要

相似文献

摘要迁移学习为将相关研究的数据整合到感兴趣的目标研究中提供了一个强大的工具。在流行病学和医学研究中,目标疾病的分类可以借用其他相关疾病和人群的信息。在这项工作中,我们考虑了高维广义线性模型(GLM)的迁移学习。提出了一种新的算法,transHDGLM,它集成了从目标研究和源研究的数据。极小极大估计的收敛速度的建立和建议的估计被证明是速率最优的。还研究了目标回归系数的统计推断。建立了一个去偏估计量的渐近正态性,该估计量可用于构造回归系数的坐标置信区间。数值研究表明,在估计和推理精度显着改善GLM,只使用目标数据。所提出的方法被应用到一个真实的数据的研究,关于使用肠道微生物组的结肠直肠癌的分类,并被证明,以提高分类精度相比,只使用目标数据的方法。本文的补充材料可在网上查阅。
Abstract Transfer learning provides a powerful tool for incorporating data from related studies into a target study of interest. In epidemiology and medical studies, the classification of a target disease could borrow information across other related diseases and populations. In this work, we consider transfer learning for high-dimensional Generalized Linear Models (GLMs). A novel algorithm, TransHDGLM, that integrates data from the target study and the source studies is proposed. Minimax rate of convergence for estimation is established and the proposed estimator is shown to be rate-optimal. Statistical inference for the target regression coefficients is also studied. Asymptotic normality for a debiased estimator is established, which can be used for constructing coordinate-wise confidence intervals of the regression coefficients. Numerical studies show significant improvement in estimation and inference accuracy over GLMs that only use the target data. The proposed methods are applied to a real data study concerning the classification of colorectal cancer using gut microbiomes, and are shown to enhance the classification accuracy in comparison to methods that only use the target data. Supplementary materials for this article are available online.