Homogeneity Structure Learning in Large-scale Panel Data with Heavy-tailed Errors

Homogeneity Structure Learning in Large-scale Panel Data with Heavy-tailed Errors
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Xiao Di;Y. Ke;Runze Li
Xiao Di;Y. Ke;Runze Li
中科院分区:
其他
文献类型:
--
作者:
Xiao Di;Y. Ke;Runze Li

文献摘要

相似文献

大规模面板数据在许多现代数据科学应用中无处不在。传统的面板数据分析方法无法解决新的挑战,如协变量的个人影响,endo-weight,嵌入式低维结构,重尾误差,所产生的创新的数据收集平台上的应用程序运行。针对这些挑战,本文采用交互效应模型研究了大规模面板数据。该模型考虑了协变量对每个空间节点的个体影响,并通过允许潜在因素影响协变量和误差来消除外生条件。此外,我们放弃了亚高斯假设,并允许误差是重尾的。此外,我们提出了一个数据驱动的程序来学习一个简约而灵活的同质性结构嵌入在高维的个人影响的协变量。同质性结构假设存在回归系数的分区,其中系数在每个组内相同,但在组之间不同。同质结构是灵活的,因为它包含许多广泛假设的低维结构(稀疏性,全局影响等)。作为其特殊情况。非渐近性质的建立,以证明所提出的学习过程。大量的数值实验证明了所提出的学习过程比传统方法的优势,特别是当数据是从重尾分布生成时。
Large-scale panel data is ubiquitous in many modern data science applications. Conventional panel data analysis methods fail to address the new challenges, like individual impacts of covariates, endogeneity, embedded low-dimensional structure, and heavy-tailed errors, arising from the innovation of data collection platforms on which applications operate. In response to these challenges, this paper studies large-scale panel data with an interactive effects model. This model takes into account the individual impacts of covariates on each spatial node and removes the exogenous condition by allowing latent factors to affect both covariates and errors. Besides, we waive the sub-Gaussian assumption and allow the errors to be heavy-tailed. Further, we propose a data-driven procedure to learn a parsimonious yet flexible homogeneity structure embedded in high-dimensional individual impacts of covariates. The homogeneity structure assumes that there exists a partition of regression coefficients where the coefficients are the same within each group but different between the groups. The homogeneity structure is flexible as it contains many widely assumed lowdimensional structures (sparsity, global impact, etc.) as its special cases. Non-asymptotic properties are established to justify the proposed learning procedure. Extensive numerical experiments demonstrate the advantage of the proposed learning procedure over conventional methods especially when the data are generated from heavy-tailed distributions.