Data reduction for prediction:: A case study on robust coding of age and family history for the risk of having a genetic mutation

Data reduction for prediction:: A case study on robust coding of age and family history for the risk of having a genetic mutation
复制标题

DOI:
10.1002/sim.3119
复制
发表时间:
2007-12-30
影响因子:
2
通讯作者:
Syngal, Sapna
Syngal, Sapna
中科院分区:
医学3区
文献类型:
--
作者:
Steyerberg, Ewout W.;Balmana, Judith;Syngal, Sapna

文献摘要

被引文献

相似文献

在预测模型的开发中通常需要减少数据,例如年龄和家族史在识别具有基因突变的受试者中的影响。我们的目的是通过相关预测变量的稳健编码来评估模型简化策略。我们考虑了 898 名疑似患有林奇综合征的患者,该综合征主要是由错配修复基因 MLH1 或 MSH2 突变引起的。通过逻辑回归分析,患者及其亲属中结直肠癌(CRC)和子宫内膜癌的存在与突变患病率相关。简化模型和更复杂模型的性能通过一致性统计量 (c) 进行量化,并通过交叉验证和引导进行乐观修正。在 1016 名患者中进行了外部验证。第一个挑战是对 CRC 诊断时的年龄进行编码,我们通过计算诊断时年龄的总和来强制患者、一级和二级亲属的效果相同。作为进一步简化,二级亲属的 CRC 诊断权重是一级亲属的一半。子宫内膜癌也遵循这些数据减少方法。对于包含单个预测变量效应的更复杂模型,简化模型使用 7 个自由度 (df),而不是 17 个自由度 (df)。乐观校正的 c 更高(0.79 而不是 0.77),但外部 c 相似(简化和更复杂的模型为 0.78)。逐步选择的模型表现稍差(外部 c=0.77)。总之,可以用相对较少的 df 开发一个预测模型,捕捉家族中每种癌症类型的患者和亲属诊断时年龄的影响。这种鲁棒编码可能尤其与相对较小的数据集中的建模相关。版权所有 (C) 2007 John Wiley & Sons, Ltd.
Data reduction is often desired in the development of a prediction model, for example for effects of age and family history in the identification of subjects having a genetic mutation. We aimed to evaluate a strategy for model simplification by robust coding of related predictors. We considered 898 patients suspected of having Lynch syndrome, which is caused primarily by mutations in the mismatch repair genes, MLH1 or MSH2. The presence of colorectal cancer (CRC) and endometrial cancer in patients and their relatives was related to mutation prevalence with logistic regression analysis. The performances of simplified and more complex models were quantified with a concordance statistic (c), which was corrected for optimism by cross-validation and bootstrapping. External validation was performed in 1016 patients.The first challenge was the coding of age at diagnosis of CRC, where we forced effects to be identical in patients, in 1st degree and in 2nd degree relatives by taking the sum of the ages at diagnosis. As a further simplification, CRC diagnosis in 2nd degree relatives was weighted half that of 1st degree relatives. These data reduction approaches were also followed for endometrial cancer. The simplified model used 7 instead of 17 degrees of freedom (df) for a more complex model incorporating individual predictor effects. The optimism-corrected c was higher (0.79 instead of 0.77), but the external c was similar (0.78 for the simplified and more complex models). A stepwise selected model performed slightly worse (external c=0.77). In conclusion, a prediction model could be developed with relatively few df that captured effects of age at diagnosis across patients and relatives per type of cancer in the family. Such robust coding may especially be relevant for modeling in relatively small data sets. Copyright (C) 2007 John Wiley & Sons, Ltd.