Imputation of missing values of tumour stage in population-based cancer registration.

Imputation of missing values of tumour stage in population-based cancer registration.
复制标题

DOI:
10.1186/1471-2288-11-129
复制
发表时间:
2011-09-19
影响因子:
4
通讯作者:
Katalinic A
Katalinic A
中科院分区:
医学3区
文献类型:
--
作者:
Eisemann N;Waldmann A;Katalinic A

文献摘要

参考文献

被引文献

相似文献

在以人群为基础的癌症登记中,缺少肿瘤分期信息的数据是一个常见的问题。如果没有适当的方法处理缺失的数据,对肿瘤分期水平的统计分析可能是有偏见的。为了确定一种有效的方法来处理肿瘤分期的缺失数据,我们用链式方程检验了不同的多重填补模型,以分析恶性黑色素瘤和女性乳腺癌病例的分期特定数量。这项分析基于德国石勒苏益格-荷尔斯泰因癌症登记处的恶性黑色素瘤数据集和女性乳腺癌数据集。根据MAR缺失模式,提取具有完整肿瘤分期信息的病例,并部分去除其分期信息,从而为每个癌症实体产生五个模拟数据集。然后用链式方程对缺失的肿瘤分期值进行多重填补,使用多分类回归、预测均值匹配、随机森林和比例抽样作为填补模型。将估计的肿瘤分期、不同阶段的病例数和多次填充后的生存曲线与观察到的结果进行比较。恶性黑色素瘤的缺失值太高,无法估计每个UICC分期的合理病例数量。然而,缺失分期的多重归因导致恶性黑色素瘤的T期病例数和乳腺癌的T期和UICC期病例数接近观察到的病例数。个体水平上观察到的肿瘤分期、不同阶段的病例数和观察到的生存曲线最好用多分类回归或预测均值匹配来满足,而不是用随机森林或比例抽样作为归因模型。这一有限的模拟研究表明,如果未分期病例的数量处于合理水平,链式方程多重归因法是处理基于人群的癌症登记中关于肿瘤分期的缺失信息的合适技术。
Missing data on tumour stage information is a common problem in population-based cancer registries. Statistical analyses on the level of tumour stage may be biased, if no adequate method for handling of missing data is applied. In order to determine a useful way to treat missing data on tumour stage, we examined different imputation models for multiple imputation with chained equations for analysing the stage-specific numbers of cases of malignant melanoma and female breast cancer. This analysis was based on the malignant melanoma data set and the female breast cancer data set of the cancer registry Schleswig-Holstein, Germany. The cases with complete tumour stage information were extracted and their stage information partly removed according to a MAR missingness-pattern, resulting in five simulated data sets for each cancer entity. The missing tumour stage values were then treated with multiple imputation with chained equations, using polytomous regression, predictive mean matching, random forests and proportional sampling as imputation models. The estimated tumour stages, stage-specific numbers of cases and survival curves after multiple imputation were compared to the observed ones. The amount of missing values for malignant melanoma was too high to estimate a reasonable number of cases for each UICC stage. However, multiple imputation of missing stage values led to stage-specific numbers of cases of T-stage for malignant melanoma as well as T- and UICC-stage for breast cancer close to the observed numbers of cases. The observed tumour stages on the individual level, the stage-specific numbers of cases and the observed survival curves were best met with polytomous regression or predictive mean matching but not with random forest or proportional sampling as imputation models. This limited simulation study indicates that multiple imputation with chained equations is an appropriate technique for dealing with missing information on tumour stage in population-based cancer registries, if the amount of unstaged cases is on a reasonable level.
DOI: 10.1016/s0959-8049(03)00259-4
发表时间: 2003-08-01
影响因子: 8.4
作者:
Duffy, SW;Tabar, L;Yen, MFA
通讯作者: Yen, MFA
DOI: 10.1093/ije/dyp309
发表时间: 2010-02-01
影响因子: 7.7
作者:
Nur, Ula;Shack, Lorraine G.;Coleman, Michel P.
通讯作者: Coleman, Michel P.
DOI: 10.1093/aje/kwq260
发表时间: 2010-11-01
影响因子: 5
作者:
Burgette, Lane F.;Reiter, Jerome P.
通讯作者: Reiter, Jerome P.
DOI: 10.1136/bmj.314.7079.472
发表时间: 1997-02-15
影响因子: --
作者:
Stockton, D;Davies, T;McCann, J
通讯作者: McCann, J
DOI: 10.1186/1471-2407-10-100
发表时间: 2010-03-16
期刊: BMC CANCER
影响因子: 3.8
作者:
Redaniel, Maria Theresa;Laudico, Adriano;Brenner, Hermann
通讯作者: Brenner, Hermann