Comparison of ordinal and nominal classification trees to predict ordinal expert-based occupational exposure estimates in a case-control study.

Comparison of ordinal and nominal classification trees to predict ordinal expert-based occupational exposure estimates in a case-control study.
复制标题

比较序数和名义分类树,以预测病例对照研究中基于专家的序数职业暴露估计。

DOI:
10.1093/annhyg/meu098
复制
发表时间:
2015
期刊:
The Annals of occupational hygiene
影响因子:
--
通讯作者:
Friesen,MelissaC
Friesen,MelissaC
中科院分区:
--
文献类型:
--
作者:
Wheeler,DavidC;Archer,KellieJ;Burstyn,Igor;Yu,Kai;Stewart,PatriciaA;Colt,JoanneS;Baris,Dalsu;Karagas,MargaretR;Schwenn,Molly;Johnson,Alison;Armenti,Karla;Silverman,DebraT;Friesen,MelissaC

文献摘要

相似文献

在病例对照研究中,为了评估职业暴露,暴露评估员通常单独审查每个工作以分配暴露估计值。这一过程缺乏透明度,也没有提供一种机制来重新创建其他研究中的决策规则。在我们以前的工作中,名义(无序分类)分类树(CT)通常成功地预测了专家评估的顺序暴露估计值(即无,低,中,高)来自职业问卷调查的答复,但仍有改进的余地。我们的目的是确定使用最近开发的有序CT是否会提高在一项病例对照研究中预测有序职业柴油机废气暴露估计的标称树的性能。MethodsWe使用了一个标称和四个有序CT方法来预测专家评估的职业柴油机废气暴露的概率、强度和频率估计(每个分类为无、低、中、高),来自新英格兰膀胱癌研究中14983份工作的问卷调查。为了复制单一树的常见用法,我们将每种方法应用于70%的工作的单一样本,使用15%进行测试,15%验证每种方法。为了表征性能的变异性,我们进行了重复采样100次的呼吸分析。我们使用索默斯d评估了树预测和专家估计之间的一致性,它测量了预测和观察得分之间的有序关联方面的差异,并且可以类似地解释为相关系数。使用二次误分类函数并基于总误分类成本控制树大小的有序CT方法具有稍微好的预测性能对于频率度量(索默斯d:名义树= 0.61;序数树= 0.63)是统计学显著的,并且对于概率(名义= 0.65;序数= 0.66)和强度(名义= 0.65;序数= 0.65)度量具有类似的性能。在所有暴露度量中,与名义树相比,最佳有序CT预测与专家评估存在较大分歧的情况较少(即,对于高暴露的工作,没有预测暴露,反之亦然)。例如,专家分配的高强度的曝光,该模型预测为没有曝光的工作的百分比为29%的名义树和22%for the best ordinal tree.ConclusionsThe整体协议是相似的CT模型,但是,使用有序模型减少了分歧发生时的差异的大小。由于性能最佳的模型可能因情况而异,研究人员应考虑评估多种CT方法,以最大限度地提高其数据的预测性能。
ObjectivesTo evaluate occupational exposures in case–control studies, exposure assessors typically review each job individually to assign exposure estimates. This process lacks transparency and does not provide a mechanism for recreating the decision rules in other studies. In our previous work, nominal (unordered categorical) classification trees (CTs) generally successfully predicted expert-assessed ordinal exposure estimates (i.e. none, low, medium, high) derived from occupational questionnaire responses, but room for improvement remained. Our objective was to determine if using recently developed ordinal CTs would improve the performance of nominal trees in predicting ordinal occupational diesel exhaust exposure estimates in a case–control study.MethodsWe used one nominal and four ordinal CT methods to predict expert-assessed probability, intensity, and frequency estimates of occupational diesel exhaust exposure (each categorized as none, low, medium, or high) derived from questionnaire responses for the 14983 jobs in the New England Bladder Cancer Study. To replicate the common use of a single tree, we applied each method to a single sample of 70% of the jobs, using 15% to test and 15% to validate each method. To characterize variability in performance, we conducted a resampling analysis that repeated the sample draws 100 times. We evaluated agreement between the tree predictions and expert estimates using Somers’d, which measures differences in terms of ordinal association between predicted and observed scores and can be interpreted similarly to a correlation coefficient.ResultsFrom the resampling analysis, compared with the nominal tree, an ordinal CT method that used a quadratic misclassification function and controlled tree size based on total misclassification cost had a slightly better predictive performance that was statistically significant for the frequency metric (Somers’d: nominal tree = 0.61; ordinal tree = 0.63) and similar performance for the probability (nominal = 0.65; ordinal = 0.66) and intensity (nominal = 0.65; ordinal = 0.65) metrics. The best ordinal CT predicted fewer cases of large disagreement with the expert assessments (i.e. no exposure predicted for a job with high exposure and vice versa) compared with the nominal tree across all of the exposure metrics. For example, the percent of jobs with expert-assigned high intensity of exposure that the model predicted as no exposure was 29% for the nominal tree and 22% for the best ordinal tree.ConclusionsThe overall agreements were similar across CT models; however, the use of ordinal models reduced the magnitude of the discrepancy when disagreements occurred. As the best performing model can vary by situation, researchers should consider evaluating multiple CT methods to maximize the predictive performance within their data.