Supervised multivariate learning with simultaneous feature auto‐grouping and dimension reduction

Supervised multivariate learning with simultaneous feature auto‐grouping and dimension reduction
复制标题

DOI:
10.1111/rssb.12492
复制
发表时间:
2021-12
期刊:
Journal of the Royal Statistical Society: Series B (Statistical Methodology)
影响因子:
--
通讯作者:
Yiyuan She;Jiahui Shen;Chao Zhang
Yiyuan She;Jiahui Shen;Chao Zhang
中科院分区:
其他
文献类型:
--
作者:
Yiyuan She;Jiahui Shen;Chao Zhang

文献摘要

相似文献

现代高维方法通常采用“稀疏性下注”原则,而在监督多元学习中,统计学家可能面临大量非零系数的“密集”问题。本文提出了一种新颖的集群降阶学习(CRL)框架,该框架采用两个联合矩阵正则化来自动对构建预测因子的特征进行分组。 CRL 比低秩建模更具可解释性,并且放宽了变量选择中严格的稀疏性假设。在本文中,提出了新的信息理论限制,以揭示寻找聚类的内在成本,以及多元学习中维数的好处。此外,还开发了一种有效的优化算法,该算法可以在保证收敛的情况下执行子空间学习和聚类。获得的定点估计量虽然不一定是全局最优的,但在某些规律性条件下具有超出标准似然设置的所需统计精度。此外,提出了一种新的信息准则及其无标度形式,用于聚类和等级选择,并且在不假设无限样本量的情况下具有严格的理论支持。广泛的模拟和真实数据实验证明了所提出方法的统计准确性和可解释性。
Modern high‐dimensional methods often adopt the ‘bet on sparsity’ principle, while in supervised multivariate learning statisticians may face ‘dense’ problems with a large number of nonzero coefficients. This paper proposes a novel clustered reduced‐rank learning (CRL) framework that imposes two joint matrix regularizations to automatically group the features in constructing predictive factors. CRL is more interpretable than low‐rank modelling and relaxes the stringent sparsity assumption in variable selection. In this paper, new information‐theoretical limits are presented to reveal the intrinsic cost of seeking for clusters, as well as the blessing from dimensionality in multivariate learning. Moreover, an efficient optimization algorithm is developed, which performs subspace learning and clustering with guaranteed convergence. The obtained fixed‐point estimators, although not necessarily globally optimal, enjoy the desired statistical accuracy beyond the standard likelihood setup under some regularity conditions. Moreover, a new kind of information criterion, as well as its scale‐free form, is proposed for cluster and rank selection, and has a rigorous theoretical support without assuming an infinite sample size. Extensive simulations and real‐data experiments demonstrate the statistical accuracy and interpretability of the proposed method.