课题基金 / 基金详情

SEER RRSS Using Multiple Imputation to Enhance the Utility of SEER Summary Stage

SEER RRSS Using Multiple Imputation to Enhance the Utility of SEER Summary Stage
SEER RRSS 使用多重插补增强 SEER 摘要阶段的实用性
批准号:
8351006
负责人:
THOMAS TUCKER
金额:
$5.3万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-30 至 2012-09-29

项目摘要

项目成果

THOMAS TUCKER的其他基金

相似基金

相关文献

中文摘要
翻译
多重填补(MI)方法已被广泛应用于许多科学领域来解决丢失数据的问题。几个统计软件包已经实施了MI程序。然而,他们的表现各不相同。尝试比较MI程序在不同统计程序包中的表现。由于MI的复杂性,基于特定的设置进行比较,例如假设缺失模式是单调缺失、缺失变量是半连续的、或者缺失数据是简单的人工数据。这些比较都不涉及在大数据集中使用哪些方法最好,例如具有非单调缺失模式和许多变量的SEER注册表数据。这项研究将调查在使用MI方法处理丢失数据时SEER注册表数据更具体的问题。通过这项研究,将为癌症登记社区的研究人员提供关于正确处理癌症登记数据中缺失数据问题的指导。 数据缺失是大多数科学研究中经常遇到的问题,也是大数据集的共同特征,尤其是医学数据集。如果处理不当,可能会导致偏差或导致分析效率低下。由于其高标准要求,SEER数据对于收集的大多数变量只有一小部分缺失数据。然而,一个非常重要的变量,SEER汇总阶段(1977、2000和CS),包含了更高比例的丢失或未知数据,特别是对于某些癌症部位。例如,在2001-2003年SEER数据的可变SEER摘要阶段2000中,9.8%的肺癌病例和22%的肝癌病例被编码为未知。 完整病例法(列表删除)是在癌症登记社区的研究人员中解决这一缺失数据问题的最常用的方法。如果丢失的数据不是完全随机丢失的,则使用完全案例方法将引入偏差并生成不正确的结果。在上述同一研究资料中,已知分期和未知期患者的癌症诊断年龄分布有显著差异:已知分期为75岁或以上者占34%,未知期为75岁或以上者占54%。这有力地表明,分期未知的病例并不是完全随机缺失的。因此,完全案例方法并不是分析数据的理想方法。将未知阶段的病例编码为一个单独的子类别肯定会包括数据分析中的所有病例,但不幸的是,当数据不是完全随机缺失的时候,这种类型的分析被发现存在严重的偏差。 与完全案例方法相比,MI是处理缺失数据的一种更复杂的方法,当数据随机缺失时,MI提供了更好的估计。MI最早于1978年提出,近年来已成为缺失数据统计分析中一种重要而有影响力的方法,因为它易于使用,并可在许多统计程序中使用。MI用一组合理的值替换每个缺失的值,这些值代表了关于最合适的值的不确定性,然后将每个完整数据集的单独数据分析的结果组合起来,以生成最终估计。有人认为,即使在假设不满足的情况下,MI通常也能提供有效和稳健的推论。
英文摘要
Multiple Imputation (MI) methods have been widely used in many scientific fields to address missing data issues. Several statistical software packages have implemented MI procedures. However, their performance varies. Attempts were made to compare performance of MI procedures in varying statistical packages. Because of the complexity of MI, comparisons were made based on specific settings, such as assuming missing patterns were monotonic missing, the missing variable was semi-continuous, or the missing data were simple artificial data. None of these comparisons addressed which methods are best used in a large data set, such as the SEER registry data with non-monotonic missing pattern and many variables. This study will investigate issues more specific to the SEER registry data when using MI methods to handle missing data. Through this study, guidance for researchers in the cancer registry community will be provided with regard to properly handling of the missing data issue in cancer registry data. Missing data is a frequent problem in most scientific studies and a common feature of large data sets in general and medical data sets in particular. It can cause bias or lead to inefficient analyses if not handled properly. Because of its high standard requirements, the SEER data have only a small fraction of missing data for most of the variables collected. However, one very important variable, the SEER Summary Stage (1977, 2000 and CS), contains a higher percentage of missing or unknown data, especially for certain cancer sites. For example, 9.8% of the lung cancer cases and 22% of the liver cancer cases were coded as unknown for the variable SEER Summary Stage 2000 for the 2001-2003 SEER data. The complete case method (listwise deletion) is the most commonly used method to address this missing data issue for data analysis among researchers in the cancer registry community. If missing data are not missing completely at random, using the complete case method will introduce bias and generate incorrect results. For the same study data mentioned above, the distributions of age at cancer diagnosis for cases with known stage and cases with unknown stage were significantly different ¿ 34% of cases were 75 years old or older for known stage while 54% of cases were 75 years old or older for unknown stage. This strongly suggests cases with unknown stage were not missing completely at random. Hence, the complete case method is not an ideal method to analyze the data. Coding the cases with unknown stage as a separate sub-category will certainly include all cases in data analysis, but unfortunately, severe bias has been found for this type of analysis when data are not missing completely at random. Compared to the complete case method, MI, one of the more sophisticated methods to handle missing data, provides superior estimates when data are missing at random. First proposed in 1978, MI has become an important and influential approach in the statistical analysis of missing data in recent years because it is easy to use and readily available in many statistical packages. MI replaces each missing value with a set of plausible values that represents the uncertainty about the most appropriate value to impute, then combines results from separate data analyses for each complete dataset to generate the final estimates. It has been suggested that MI often provides valid and robust inferences even when assumptions were not met.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SEER-LINKED VIRTUAL TISSUE REPOSITORY (VTR) PROGRAM
  • 批准号:
    10976186
  • 项目类别:
  • 资助金额:
    $14.28万
  • 财政年份:
    2023
  • 负责人:
    THOMAS TUCKER
  • 依托单位:
NCI SEER-LINKED PEDIATRIC WHOLE SLIDE IMAGING (POP: 8/17/2020 - 8/16/2021)
  • 批准号:
    10272821
  • 项目类别:
  • 资助金额:
    $20.36万
  • 财政年份:
    2020
  • 负责人:
    THOMAS TUCKER
  • 依托单位:
Patterns of Care/Quality of Care Study: Diagnosis Year 2013 (SEER)Period of Performance: 08/15/2014-08/14/2015Line item #: 1
  • 批准号:
    8928277
  • 项目类别:
  • 资助金额:
    $11.67万
  • 财政年份:
    2014
  • 负责人:
    THOMAS TUCKER
  • 依托单位:
IGF::OT::IGF 402 - NORTH AMERICAN ASSOCIATION OF CENTRAL CANCER REGISTRIES, INC. (NAACCR); TECHNICAL SUPPORT FOR CANCER SURVEILLANCE; POP 07/01/2014-06/30/2014
  • 批准号:
    8885293
  • 项目类别:
  • 资助金额:
    $154.61万
  • 财政年份:
    2014
  • 负责人:
    THOMAS TUCKER
  • 依托单位:
海外基金