Bayesian profile regression with an application to the National survey of children's health

Bayesian profile regression with an application to the National survey of children's health
复制标题

DOI:
10.1093/biostatistics/kxq013
复制
发表时间:
2010-07-01
期刊:
影响因子:
2.1
通讯作者:
Richardson, Sylvia
Richardson, Sylvia
中科院分区:
数学2区
文献类型:
--
作者:
Molitor, John;Papathomas, Michail;Richardson, Sylvia

文献摘要

被引文献

相似文献

当试图使用包含数十个潜在相关变量的数据集进行超出主效应的推断时,标准回归分析经常会遇到问题。这种情况会出现,例如,在流行病学中,由大量问题组成的调查或研究问卷产生一组潜在的笨拙的相互关联的数据,从中梳理出多个协变量的影响是困难的。我们提出了一种方法,通过使用由协变量值序列形成的轮廓作为其基本推理单位,来解决分类协变量的这些问题。这些协变量概况被聚类成组,并通过回归模型与相关结果相关联。所提出的建模框架的贝叶斯聚类方面与传统聚类方法相比具有许多优点,因为它允许组的数量变化,揭示子组并检查它们与感兴趣的结果的关联,并将模型拟合为一个单元,允许个人的结果可能影响集群成员。通过对全国儿童健康调查获得的调查数据的分析,证明了该方法。该方法是使用标准的贝叶斯建模软件WinBUGS实现的,其代码可在Biostatistics在线上的补充材料中获得。此外,我们开发的一些后处理工具可以帮助解释数据分区。
Standard regression analyses are often plagued with problems encountered when one tries to make inference going beyond main effects using data sets that contain dozens of variables that are potentially correlated. This situation arises, for example, in epidemiology where surveys or study questionnaires consisting of a large number of questions yield a potentially unwieldy set of interrelated data from which teasing out the effect of multiple covariates is difficult. We propose a method that addresses these problems for categorical covariates by using, as its basic unit of inference, a profile formed from a sequence of covariate values. These covariate profiles are clustered into groups and associated via a regression model to a relevant outcome. The Bayesian clustering aspect of the proposed modeling framework has a number of advantages over traditional clustering approaches in that it allows the number of groups to vary, uncovers subgroups and examines their association with an outcome of interest, and fits the model as a unit, allowing an individual's outcome potentially to influence cluster membership. The method is demonstrated with an analysis of survey data obtained from the National Survey of Children's Health. The approach has been implemented using the standard Bayesian modeling software, WinBUGS, with code provided in the supplementary material available at Biostatistics online. Further, interpretation of partitions of the data is helped by a number of postprocessing tools that we have developed.