Statistical significance in high-dimensional linear mixed models.

Statistical significance in high-dimensional linear mixed models.
复制标题

DOI:
10.1145/3412815.3416883
复制
发表时间:
2020-10
期刊:
FODS '20 : proceedings of the 2020 ACM-IMS Foundations of Data Science Conference : October 19-20, 2020, Virtual Event, USA. ACM-IMS Foundations of Data Science Conference (2020 : Online)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

本文讨论了高维线性混合效应模型的推理框架。例如,当我们对M个受试者进行n次重复测量时,这些是合适的模型。我们考虑一个场景,其中固定效应的数量p很大(可能大于M),但随机效应的数量q很小。我们的框架受到了最近一系列工作的启发,这些工作提出了去偏惩罚估计量,以仅对具有固定效应的高维线性模型进行推断。特别是,我们演示了如何正确的“天真”岭估计的扩展工作,建立渐近有效的置信区间的混合效应模型。我们验证了我们的理论结果与数值实验中,我们表明我们的方法优于那些未能考虑随机效应引起的相关性。对于一个实际的演示,我们认为核黄素生产数据集,具有组结构,并表明,使用我们的方法得出的结论是一致的,在类似的数据集没有组结构。
This paper concerns the development of an inferential framework for high-dimensional linear mixed effect models. These are suitable models, for instance, when we have n repeated measurements for M subjects. We consider a scenario where the number of fixed effects p is large (and may be larger than M), but the number of random effects q is small. Our framework is inspired by a recent line of work that proposes de-biasing penalized estimators to perform inference for high-dimensional linear models with fixed effects only. In particular, we demonstrate how to correct a ‘naive’ ridge estimator in extension of work by to build asymptotically valid confidence intervals for mixed effect models. We validate our theoretical results with numerical experiments, in which we show our method outperforms those that fail to account for correlation induced by the random effects. For a practical demonstration we consider a riboflavin production dataset that exhibits group structure, and show that conclusions drawn using our method are consistent with those obtained on a similar dataset without group structure.