Classification and regression trees for epidemiologic research: an air pollution example.

Classification and regression trees for epidemiologic research: an air pollution example.
复制标题

流行病学研究的分类和回归树:空气污染的例子。

DOI:
10.1186/1476-069x-13-17
复制
发表时间:
2014-03-13
期刊:
Environmental health : a global access science source
影响因子:
--
通讯作者:
Strickland MJ
Strickland MJ
中科院分区:
其他
文献类型:
--
作者:
Gass K;Klein M;Chang HH;Flanders WD;Strickland MJ

文献摘要

被引文献

相似文献

确定和描述混合暴露如何与健康终点相关联是一项挑战。我们演示了如何分类和回归树可以用来产生假设的联合效应,从曝光混合物。我们通过调查格鲁吉亚亚特兰大市儿童哮喘急诊科就诊时CO、NO2、O3和PM2.5的联合影响来说明这种方法。污染物浓度按四分位数分类。所有污染物处于最低四分位数的天数作为参考组(n = 131),其余3,879天用于估计回归树。污染物被参数化为代表四分位数的每个顺序分割的二分变量(例如,比较CO四分位数1与CO四分位数2-4),并在泊松病例交叉模型中一次考虑一个,并控制混杂因素。选择产生最小P值的污染物-分裂作为第一分裂,并相应地划分数据集。对每个数据子集重复该过程,直到剩余分割的P值不低于给定的α,从而形成“终端节点”。我们使用病例交叉模型来估计每个终端节点与参考组相比的调整后风险比,以及在最终模型中包含终端节点的似然比检验。最大风险比对应于PM2.5处于最高四分位数而NO2处于最低两个四分位数的天数(RR:1.10,95%CI:1.05,1.16)。模型中包含所有终末节点的同步Wald检验具有显著性,卡方统计量为34.3(p = 0.001,自由度为13)。回归树可用于假设暴露混合物的联合效应,并且在空气污染流行病学领域可能特别有用,以便更好地了解复杂的多污染物暴露。
Identifying and characterizing how mixtures of exposures are associated with health endpoints is challenging. We demonstrate how classification and regression trees can be used to generate hypotheses regarding joint effects from exposure mixtures. We illustrate the approach by investigating the joint effects of CO, NO2, O3, and PM2.5 on emergency department visits for pediatric asthma in Atlanta, Georgia. Pollutant concentrations were categorized as quartiles. Days when all pollutants were in the lowest quartile were held out as the referent group (n = 131) and the remaining 3,879 days were used to estimate the regression tree. Pollutants were parameterized as dichotomous variables representing each ordinal split of the quartiles (e.g. comparing CO quartile 1 vs. CO quartiles 2–4) and considered one at a time in a Poisson case-crossover model with control for confounding. The pollutant-split resulting in the smallest P-value was selected as the first split and the dataset was partitioned accordingly. This process repeated for each subset of the data until the P-values for the remaining splits were not below a given alpha, resulting in the formation of a “terminal node”. We used the case-crossover model to estimate the adjusted risk ratio for each terminal node compared to the referent group, as well as the likelihood ratio test for the inclusion of the terminal nodes in the final model. The largest risk ratio corresponded to days when PM2.5 was in the highest quartile and NO2 was in the lowest two quartiles (RR: 1.10, 95% CI: 1.05, 1.16). A simultaneous Wald test for the inclusion of all terminal nodes in the model was significant, with a chi-square statistic of 34.3 (p = 0.001, with 13 degrees of freedom). Regression trees can be used to hypothesize about joint effects of exposure mixtures and may be particularly useful in the field of air pollution epidemiology for gaining a better understanding of complex multipollutant exposures.