Robust distributed lag models using data adaptive shrinkage.

Robust distributed lag models using data adaptive shrinkage.
复制标题

使用数据自适应收缩的鲁棒分布式滞后模型。

DOI:
10.1093/biostatistics/kxx041
复制
发表时间:
2018
期刊:
Biostatistics (Oxford, England)
影响因子:
--
通讯作者:
Coull,BrentA
Coull,BrentA
中科院分区:
--
文献类型:
--
作者:
Chen,Yin-Hsiu;Mukherjee,Bhramar;Adar,SaraD;Berrocal,VeronicaJ;Coull,BrentA

文献摘要

相似文献

分布滞后模型(DLM)已被广泛用于环境流行病学,以量化空气污染对死亡率或心血管事件等相关结果的滞后影响。一般来说,DLMs可以应用于时间序列数据,其中自变量的当前度量及其滞后度量共同影响因变量的当前度量。相应的分布滞后(DL)函数表示滞后与滞后暴露变量的系数之间的关系。常见的选择包括多项式和样条。一方面,这种约束DLM将系数指定为滞后的函数,并减少了要估计的参数的数量;因此,可以实现更高的效率。另一方面,在违反关于DL函数的假设的情况下,效应估计可能严重偏倚。在这篇文章中,我们提出了一个通用的框架收缩系数估计从无约束的DLM,这是无偏的,但可能是低效的,对系数估计从约束的DLM,以实现偏差方差权衡。收缩量可以通过多种方式确定,我们探索了几种这样的方法:经验贝叶斯型收缩,分层贝叶斯方法和广义岭回归。我们还考虑了一个两阶段收缩方法,该方法随着滞后的增加,强制效应估计值接近零。我们通过广泛的模拟研究对比了各种方法,并表明收缩方法在不同情景下的均方误差(MSE)方面具有更好的平均性能。我们使用国家死亡率,死亡率和空气污染研究(NMMAPS)的数据来说明这些方法,以探索芝加哥,IL,从1987年到2000年。
Distributed lag models (DLMs) have been widely used in environmental epidemiology to quantify the lagged effects of air pollution on an outcome of interest such as mortality or cardiovascular events. Generally speaking, DLMs can be applied to time-series data where the current measure of an independent variable and its lagged measures collectively affect the current measure of a dependent variable. The corresponding distributed lag (DL) function represents the relationship between the lags and the coefficients of the lagged exposure variables. Common choices include polynomials and splines. On one hand, such a constrained DLM specifies the coefficients as a function of lags and reduces the number of parameters to be estimated; hence, higher efficiency can be achieved. On the other hand, under violation of the assumption about the DL function, effect estimates can be severely biased. In this article, we propose a general framework for shrinking coefficient estimates from an unconstrained DLM, that are unbiased but potentially inefficient, toward the coefficient estimates from a constrained DLM to achieve a bias-variance trade-off. The amount of shrinkage can be determined in various ways, and we explore several such methods: empirical Bayes-type shrinkage, a hierarchical Bayes approach, and generalized ridge regression. We also consider a two-stage shrinkage approach that enforces the effect estimates to approach zero as lags increase. We contrast the various methods via an extensive simulation study and show that the shrinkage methods have better average performance across different scenarios in terms of mean squared error (MSE).We illustrate the methods by using data from the National Morbidity, Mortality, and Air Pollution Study (NMMAPS) to explore the association between PM, O, and SOon three types of disease event counts in Chicago, IL, from 1987 to 2000.