Sequential Estimation for Mixture of Regression Models for Heterogeneous Population

Sequential Estimation for Mixture of Regression Models for Heterogeneous Population
复制标题

DOI:
10.1016/j.csda.2024.107942
复制
发表时间:
2024-02
期刊:
Computational Statistics & Data Analysis
影响因子:
--
通讯作者:
Na You;Hongsheng Dai;Xueqin Wang;Qingyun Yu
Na You;Hongsheng Dai;Xueqin Wang;Qingyun Yu
中科院分区:
其他
文献类型:
--
作者:
Na You;Hongsheng Dai;Xueqin Wang;Qingyun Yu

文献摘要

相似文献

患者间异质性在临床研究中普遍存在,给医学研究带来了挑战。人们普遍认为,种群中存在各种亚型,它们彼此不同。识别亚型从而定制疾病预防和治疗的方法被称为精准医学。混合模型是将异质总体聚类为同质子总体的经典统计模型。然而,对于具有多个组成部分的高度异质的总体,由于EM算法对初始值的依赖性,其参数估计和聚类结果可能是模糊的。对于子分型的目的,有限的混合物的回归模型与伴随变量被认为是一种新的统计方法,提出了确定的主要成分的混合物中的大比例顺序。与已有的典型统计推断方法相比,新方法不仅不需要预先指定模型拟合的分量个数,而且参数估计和聚类结果更加可靠。仿真研究表明了该方法的优越性。对药物反应预测的真实的数据分析表明其参数估计的可靠性和识别重要亚组的能力。
Heterogeneity among patients commonly exists in clinical studies and leads to challenges in medical research. It is widely accepted that there exist various sub-types in the population and they are distinct from each other. The approach of identifying the sub-types and thus tailoring disease prevention and treatment is known as precision medicine. The mixture model is a classical statistical model to cluster the heterogeneous population into homogeneous sub-populations. However, for the highly heterogeneous population with multiple components, its parameter estimation and clustering results may be ambiguous due to the dependence of the EM algorithm on the initial values. For sub-typing purposes, the finite mixture of regression models with concomitant variables is considered and a novel statistical method is proposed to identify the main components with large proportions in the mixture sequentially. Compared to existing typical statistical inferences, the new method not only requires no pre-specification on the number of components for model fitting, but also provides more reliable parameter estimation and clustering results. Simulation studies demonstrated the superiority of the proposed method. Real data analysis on the drug response prediction illustrated its reliability in the parameter estimation and capability to identify the important subgroup.