Adverse subpopulation regression for multivariate outcomes with high-dimensional predictors.

Adverse subpopulation regression for multivariate outcomes with high-dimensional predictors.
复制标题

具有高维预测变量的多变量结果的不利亚群回归。

DOI:
10.1002/sim.5520
复制
发表时间:
2012
影响因子:
2
通讯作者:
Ashley-Koch,AllisonE
Ashley-Koch,AllisonE
中科院分区:
医学3区
文献类型:
--
作者:
Zhu,Bin;Dunson,DavidB;Ashley-Koch,AllisonE

文献摘要

相似文献

生物医学研究在评估多个相关健康结果和高维预测因子之间的关系方面有着共同的兴趣。例如,在生殖流行病学中,人们可以收集妊娠结果,如妊娠期和出生体重,以及预测因子,如多个候选基因中的单核苷酸多态性和环境暴露。在这种情况下,需要一种简单而灵活的方法来从一组高维候选预测因子中选择不良健康反应的真实预测因子。为了解决这个问题,可以考虑连续结果的线性回归模型,或者使用预定义的截止值将这些结果转换为不良反应的二元指标。前一种策略的缺点是往往会导致拟合不佳的模型,不能很好地预测风险,而后一种方法可能对临界值的选择非常敏感。作为一种简单而灵活的替代方案,我们提出了一种用于不利亚群回归的方法,该方法依赖于双组分潜在类模型,其中主导组分对应于(假定的)健康个体,而落入少数组分的风险通过逻辑回归来表征。逻辑回归模型旨在通过使用灵活的非参数多重收缩方法来适应高维预测因子,如在具有大量基因与环境相互作用的研究中所发生的。吉布斯采样器是为后验计算而开发的。我们评估的方法与模拟研究的使用,并将其应用于遗传流行病学研究的妊娠结局。版权所有© 2012约翰威利父子有限公司.
Biomedical studies have a common interest in assessing relationships between multiple related health outcomes and high‐dimensional predictors. For example, in reproductive epidemiology, one may collect pregnancy outcomes such as length of gestation and birth weight and predictors such as single nucleotide polymorphisms in multiple candidate genes and environmental exposures. In such settings, there is a need for simple yet flexible methods for selecting true predictors of adverse health responses from a high‐dimensional set of candidate predictors. To address this problem, one may either consider linear regression models for the continuous outcomes or convert these outcomes into binary indicators of adverse responses using predefined cutoffs. The former strategy has the disadvantage of often leading to a poorly fitting model that does not predict risk well, whereas the latter approach can be very sensitive to the cutoff choice. As a simple yet flexible alternative, we propose a method for adverse subpopulation regression, which relies on a two‐component latent class model, with the dominant component corresponding to (presumed) healthy individuals and the risk of falling in the minority component characterized via a logistic regression. The logistic regression model is designed to accommodate high‐dimensional predictors, as occur in studies with a large number of gene by environment interactions, through the use of a flexible nonparametric multiple shrinkage approach. The Gibbs sampler is developed for posterior computation. We evaluate the methods with the use of simulation studies and apply these to a genetic epidemiology study of pregnancy outcomes. Copyright © 2012 John Wiley & Sons, Ltd.