课题基金 / 基金详情

Semiparametric Estimation and Variable Selection in the Presence of Nonignorable Nonresponse

Semiparametric Estimation and Variable Selection in the Presence of Nonignorable Nonresponse
存在不可忽略的无反应时的半参数估计和变量选择
批准号:
1612873
负责人:
Jun Shao
金额:
$29.04万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2020-08-31

项目摘要

项目成果

Jun Shao的其他基金

相似基金

相关文献

中文摘要
翻译
无应答现象存在于许多统计应用中。在大多数调查问题中,许多抽样单位未能提供部分或全部调查问题的答案。 在医学或健康研究中,不完整数据的百分比往往是可观的。当无应答与缺失数据相关时,处理无应答非常具有挑战性。由于这项研究的动机是调查机构,如美国人口普查局和加拿大统计局的问题,或在医疗和健康研究的数据集,从这项研究中获得的结果将有显着的影响,在实践中的估计和推断处理无应答的方法。这项研究的结果也将为这一领域的进一步研究提供启示。当无应答机制或倾向仅取决于观察到的数据时,无应答被称为可验证的;否则,它是不可验证的。有丰富的文献方法处理可排除的无应答。处理不可验证的无应答更具挑战性,因为必须施加假设以确保未知人群特征的可识别性和可估计性,并且这些假设很难使用不可验证的无应答数据进行检查。将为可解释的无应答而开发的方法应用于具有不可解释的无应答的数据可能会在统计估计和推断中产生严重的偏差。本研究的重点是估计的基础上的数据与不可否认的无响应在以下两个一般性的主题。(1)半参数估计如果对无反应倾向和关注的群体分布假设一个完全参数模型,那么在某些可识别性假设下,可以使用参数似然来推导关注参数的有效估计。 然而,这种参数化方法对模型误设很敏感,特别是当无响应是不可解释的时候。另一方面,与无无应答的情况不同,纯粹的非参数方法无法识别总体。本研究研究半参数方法,假设倾向或总体分布的一个成分是参数的,而其他成分是非参数的。将努力研究在不同假设下各种方法的稳健性和有效性,纵向或多变量结局与不可解释的无应答,缺失结局和协变量的问题,以及Meta分析中未测量的混杂因素或系统性缺失协变量数据。(2)模型和变量选择。当无应答不可验证时,需要使用称为无应答工具的协变量,该协变量始终可以观察到,并有助于识别群体参数。此外,必须假设倾向或人口分布的参数组成部分。因此,需要进行模型和/或变量选择,以确保假设的参数分量和选择的无响应工具是适当的。由于不可解释的无响应,现有的模型和变量选择技术是不适用的。这项研究将为模型选择和无响应仪器的选择开发新的技术。此外,在大数据时代,存在一个非常大的辅助变量集,可以用作协变量,本研究将研究降维和变量选择,以准确估计存在不可解释的无响应。
英文摘要
Nonresponse exists in many statistical applications. In most survey problems, many sampled units fail to provide answers to some or all survey questions. In medical or health studies, the percentages of incomplete data are often appreciable. Handling nonresponse is very challenging when nonresponse is related to the missing data. Since this research is motivated by problems in survey agencies such as the U.S. Census Bureau and Statistics Canada, or by data sets in medical and health studies, results obtained from this research will have significant impacts on the methodology of handling nonresponse for estimation and inference in practice. The results from this research will also shed light on further research in this area. When the nonresponse mechanism or propensity depends on observed data only, the nonresponse is called ignorable; otherwise, it is nonignorable. There is a rich literature on methodology of handling ignorable nonresponse. Handling nonignorable nonresponse is much more challenging, since assumptions have to be imposed to ensure the identifiability and estimability of unknown population characteristics and these assumptions are hard to check using data with nonignorable nonresponse. Applying methods developed for ignorable nonresponse to data with nonignorable nonresponse may create serious biases in statistical estimation and inference. This research focuses on estimation based on data with nonignorable nonresponse in the following two general topics. (1) Semiparametric estimation. If a fully parametric model is assumed on the nonresponse propensity and the population distribution of interest, then valid estimators of parameters of interest may be derived using the parametric likelihood under some identifiability assumption. However, this parametric approach is sensitive to model misspecification, especially when nonresponse is nonignorable. On the other hand, unlike the situation with no nonresponse, a purely nonparametric approach cannot identify the population. This research studies semiparametric methods, assuming one component of the propensity or population distribution is parametric and the others are nonparametric. Efforts will be made to study robustness and efficiency of various methods under different assumptions, longitudinal or multivariate outcomes with nonignorable nonresponse, problems with both missing outcomes and covariates, and unmeasured confounders or systematic missing covariate data in meta analyses. (2) Model and variable selection. When nonresponse is nonignorable, a covariate called nonresponse instrument needs to be used, which is always observed and helps to identify population parameters. In addition, a parametric component of either the propensity or the population distribution has to be assumed. Thus, it is desired to perform model and/or variable selection to ensure that the assumed parametric component and the selected nonresponse instrument are appropriate. Because of nonignorable nonresponse, the existing model and variable selection techniques are not applicable. This research will develop new techniques for model selection and the selection of nonresponse instruments. Furthermore, in the big data era there exists an extremely large set of auxiliary variables that can be used as covariates and this research will study dimension reduction and variable selection for accurate estimation in the presence of nonignorable nonresponse.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Variable Selection, Instrument Search and Estimation in Problems with Nonignorable Missing Data
  • 批准号:
    1914411
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2019
  • 负责人:
    Jun Shao
  • 依托单位:
Analysis of Longitudinal or Multivariate Data with Nonignorable Missing Values
  • 批准号:
    1305474
  • 项目类别:
    Standard Grant
  • 资助金额:
    $18.0万
  • 财政年份:
    2013
  • 负责人:
    Jun Shao
  • 依托单位:
Inference with Survey Data Having Nonignorable Nonresponse
  • 批准号:
    1007454
  • 项目类别:
    Standard Grant
  • 资助金额:
    $22.16万
  • 财政年份:
    2010
  • 负责人:
    Jun Shao
  • 依托单位:
Analysis of Survey Data Using Imputation for Nonrespondents
  • 批准号:
    0705033
  • 项目类别:
    Standard Grant
  • 资助金额:
    $21.64万
  • 财政年份:
    2007
  • 负责人:
    Jun Shao
  • 依托单位:
海外基金