Variational Bayes Ensemble Learning Neural Networks With Compressed Feature Space

Variational Bayes Ensemble Learning Neural Networks With Compressed Feature Space
复制标题

DOI:
10.1109/tnnls.2022.3172276
复制
发表时间:
2022-05
影响因子:
10.4
通讯作者:
Zihuan Liu;Shrijita Bhattacharya;T. Maiti
Zihuan Liu;Shrijita Bhattacharya;T. Maiti
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zihuan Liu;Shrijita Bhattacharya;T. Maiti

文献摘要

相似文献

我们考虑高维输入向量的非参数分类问题(小n大p问题)。为了处理高维特征空间,我们提出了一个随机投影(RP)的特征空间,然后在压缩的特征空间上训练神经网络(NN)。与正则化技术(套索,脊等)不同,基于压缩特征空间的神经网络在全数据上训练,具有显著更低的计算复杂度和存储器存储要求。尽管如此,基于随机压缩的方法通常对压缩的选择敏感。为了解决这个问题,我们采用贝叶斯模型平均(BMA)方法,并利用后验模型权重来确定:1)每次压缩下的不确定性和2)特征空间的内在维度(用于预测的特征空间的有效维度)。通过对具有接近内在维度的投影维度的模型进行平均,最终预测得到改进。此外,我们提出了一个变分的方法,上述BMA允许同时估计模型权重和模型特定的参数。由于所提出的变分解决方案是跨压缩并行,它保留了频率主义合奏技术的计算增益,同时提供了完整的不确定性量化的贝叶斯方法。我们建立了所提出的算法下的RP和先验参数的适当表征的渐近一致性。最后,我们提供了大量的数值例子,所提出的方法的实证验证。
We consider the problem of nonparametric classification from a high-dimensional input vector (small $n$ large $p$ problem). To handle the high-dimensional feature space, we propose a random projection (RP) of the feature space followed by training of a neural network (NN) on the compressed feature space. Unlike regularization techniques (lasso, ridge, etc.), which train on the full data, NNs based on compressed feature space have significantly lower computation complexity and memory storage requirements. Nonetheless, a random compression-based method is often sensitive to the choice of compression. To address this issue, we adopt a Bayesian model averaging (BMA) approach and leverage the posterior model weights to determine: 1) uncertainty under each compression and 2) intrinsic dimensionality of the feature space (the effective dimension of feature space useful for prediction). The final prediction is improved by averaging models with projected dimensions close to the intrinsic dimensionality. Furthermore, we propose a variational approach to the afore-mentioned BMA to allow for simultaneous estimation of both model weights and model-specific parameters. Since the proposed variational solution is parallelizable across compressions, it preserves the computational gain of frequentist ensemble techniques while providing the full uncertainty quantification of a Bayesian approach. We establish the asymptotic consistency of the proposed algorithm under the suitable characterization of the RPs and the prior parameters. Finally, we provide extensive numerical examples for empirical validation of the proposed method.