Shrinkage estimators for covariance matrices

Shrinkage estimators for covariance matrices
复制标题

DOI:
10.1111/j.0006-341x.2001.01173.x
复制
发表时间:
2001-12-01
期刊:
影响因子:
1.9
通讯作者:
Kass, RE
Kass, RE
中科院分区:
数学3区
文献类型:
--
作者:
Daniels, MJ;Kass, RE

文献摘要

被引文献

相似文献

许多作者已经研究了小样本中协方差矩阵的估计。标准估计器,例如非结构化最大似然估计器 (ML) 或受限最大似然 (REML) 估计器,可能非常不稳定,最小估计特征值太小,最大估计特征值太大。在小样本中更稳定地估计矩阵的标准方法是在一些简单的结构下计算 ML 或 REML 估计器,该结构涉及较少的参数估计,例如复合对称性或独立性。然而,除非假设的结构正确,否则这些估计量将不一致。如果兴趣集中于使用相关(或纵向)数据估计回归系数,则可以使用协方差矩阵的夹心估计器来提供估计系数的标准误差,该估计系数在协方差结构的错误指定下保持一致的意义上是鲁棒的。然而,对于大型矩阵,夹心估计器的低效率变得令人担忧。我们在这里考虑两种通用的收缩方法来估计协方差矩阵和回归系数。第一个涉及缩小非结构化 ML 或 REML 估计器的特征值。第二个涉及将非结构化估计器缩小为结构化估计器。对于这两种情况,数据决定收缩量。这些估计量是一致的,并对回归系数给出一致且渐近有效的估计。仿真显示了有限样本中协方差矩阵和回归系数的收缩估计器的改进操作特性。最终选择的估计器包括两种收缩方法的组合,即收缩特征值,然后收缩到结构。我们说明了我们的睡眠脑电图研究方法,该研究需要估计 24 x 24 协方差矩阵,并且对平均参数的推断关键取决于所选的协方差估计器。我们建议使用特定的收缩估计器进行推断,该估计器在结构化和非结构化估计器之间提供合理的折衷。
Estimation of covariance matrices in small samples has been studied by many authors. Standard estimators, like the unstructured maximum likelihood estimator (ML) or restricted maximum likelihood (REML) estimator, can be very unstable with the smallest estimated eigenvalues being too small and the largest too big. A standard approach to more stably estimating the matrix in small samples is to compute the ML or REML estimator under some simple structure that involves estimation of fewer parameters, such as compound symmetry or independence. However, these estimators will not be consistent unless the hypothesized structure is correct. If interest focuses on estimation of regression coefficients with correlated (or longitudinal) data, a sandwich estimator of the covariance matrix may be used to provide standard errors for the estimated coefficients that are robust in the sense that they remain consistent under misspecification of the covariance structure. With large matrices, however, the inefficiency of the sandwich estimator becomes worrisome. We consider here two general shrinkage approaches to estimating the covariance matrix and regression coefficients. The first involves shrinking the eigenvalues of the unstructured ML or REML estimator. The second involves shrinking an unstructured estimator toward a structured estimator. For both cases, the data determine the amount of shrinkage. These estimators are consistent and give consistent and asymptotically efficient estimates for regression coefficients. Simulations show the improved operating characteristics of the shrinkage estimators of the covariance matrix and the regression coefficients in finite samples. The final estimator chosen includes a combination of both shrinkage approaches, i.e., shrinking the eigenvalues and then shrinking toward structure. We illustrate our approach on a sleep EEG study that requires estimation of a 24 x 24 covariance matrix and for which inferences on mean parameters critically depend on the covariance estimator chosen. We recommend making inference using a particular shrinkage estimator that provides a reasonable compromise between structured and unstructured estimators.