Joint mean–covariance estimation via the horseshoe

Joint mean–covariance estimation via the horseshoe
复制标题

通过马蹄形进行联合均值协方差估计

DOI:
10.1016/j.jmva.2020.104716
复制
发表时间:
2021
影响因子:
1.6
通讯作者:
Bhadra, Anindya
Bhadra, Anindya
中科院分区:
数学2区
文献类型:
--
作者:
Li, Yunfan;Datta, Jyotishka;Craig, Bruce A.;Bhadra, Anindya

文献摘要

相似文献

看似无关回归是回归多个预测因子的多个相关响应的自然框架。该模型非常灵活,多元线性回归和协方差选择模型是特例。然而,它在贝叶斯框架下的基因组数据分析的实际部署是有限的,由于统计和计算的挑战。统计学上的挑战是,人们需要推断均值向量和逆协方差矩阵,这是一个比单独估计每一个更复杂的问题。计算挑战是由于参数空间的维数通常超过样本量。我们建议使用马蹄先验的均值向量和逆协方差矩阵。在分别估计均值向量或逆协方差矩阵时,该先验已证明具有出色的性能。目前的工作表明,这些优点也同时解决。提出了一个完整的贝叶斯处理,与抽样算法,是线性的预测。实现该算法的MATLAB代码可从github https://github.com/liyf1988/HS_GHS免费获得。广泛的性能比较提供了频率论和贝叶斯替代品,估计和预测性能的基因组数据集上进行了验证。
Seemingly unrelated regression is a natural framework for regressing multiple correlated responses on multiple predictors. The model is very flexible, with multiple linear regression and covariance selection models being special cases. However, its practical deployment in genomic data analysis under a Bayesian framework is limited due to both statistical and computational challenges. The statistical challenge is that one needs to infer both the mean vector and the inverse covariance matrix, a problem inherently more complex than separately estimating each. The computational challenge is due to the dimensionality of the parameter space that routinely exceeds the sample size. We propose the use of horseshoe priors on both the mean vector and the inverse covariance matrix. This prior has demonstrated excellent performance when estimating a mean vector or inverse covariance matrix separately. The current work shows these advantages are also present when addressing both simultaneously. A full Bayesian treatment is proposed, with a sampling algorithm that is linear in the number of predictors. MATLAB code implementing the algorithm is freely available from github at https://github.com/liyf1988/HS_GHS. Extensive performance comparisons are provided with both frequentist and Bayesian alternatives, and both estimation and prediction performances are verified on a genomic data set.