Variable selection in distributed sparse regression under memory constraints

Variable selection in distributed sparse regression under memory constraints
复制标题

内存约束下分布式稀疏回归的变量选择

DOI:
--
复制
发表时间:
2023
影响因子:
0.9
通讯作者:
Jiang Jiancheng
Jiang Jiancheng
中科院分区:
数学4区
文献类型:
--
作者:
Wang Haofeng;Jiang Xuejun;Zhou Min;Jiang Jiancheng

文献摘要

相似文献

本文研究了有限记忆约束下大样本分布稀疏回归的惩罚似然变量选择问题。这是大数据时代急需解决的研究问题。解决这个问题的一种简单的分治方法是将整个数据分成N个部分,并在N台机器中的一台机器上运行每个部分,通过平均来聚合所有机器的结果,最终获得选定的变量。然而,它倾向于选择更多的噪声变量,并且错误发现率可能不能很好地控制。在聚合时,我们通过特殊设计的加权平均来改进它。虽然交替方向乘子法(ADMM)可以用来处理大量的数据在文献中,我们提出的方法减少了计算负担很多,并在大多数情况下表现出更好的均方误差。在理论上,我们建立了参数个数不同的似然模型的估计量的渐近性质。在一定的规律性条件下,我们建立了oracle.properties,在这个意义上,我们的分布式估计共享相同的渐近效率的估计的基础上的完整的样本。在计算上,提出了一种分布式惩罚似然算法,以改进一般似然的上下文中的结果。最后,通过仿真和一个真实的算例对所提出的方法进行了评价.
This paper studies variable selection using the penalized likelihood method for distributed sparse regression with large sample size n under a limited memory constraint. This is a much needed research problem to be solved in the big data era. A naive divide-and-conquer method solving this problem is to split the whole data into N parts and run each part on one of N machines, aggregate the results from all machines via averaging, and finally obtain the selected variables. However, it tends to select more noise variables, and the false discovery rate may not be well controlled. We improve it by a special designed weighted average in aggregation. Although the alternating direction method of multiplier (ADMM) can be used to deal with massive data in the literature, our proposed method reduces the computational burden a lot and performs better by mean square error in most cases. Theoretically, we establish asymptotic properties of the resulting estimators for the likelihood models with a diverging number of parameters. Under some regularity conditions we establish oracle.properties in the sense that our distributed estimator shares the same asymptotic efficiency as the estimator based on the full sample. Computationally, a distributed penalized likelihood algorithm is proposed to refine the results in the context of general likelihoods. Furthermore, the proposed.method is evaluated by simulations and a real example.