课题基金 / 基金详情

Semiparametric Efficient Estimation of Models of Measurement Errors and Missing Data

Semiparametric Efficient Estimation of Models of Measurement Errors and Missing Data
测量误差和缺失数据模型的半参数高效估计
批准号:
0452143
负责人:
Han Hong
金额:
$11.63万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-04-15 至 2007-03-31

项目摘要

项目成果

Han Hong的其他基金

相似基金

相关文献

中文摘要
翻译
经济学中的许多经验研究因相关变量的存在而变得复杂,这些变量没有被观察到,要么是因为它们只以不完整或腐败的方式获得,要么是因为它们本身就无法观察到。重要的例子包括面板数据分析中的损耗和普遍存在的测量误差,这些误差可能与真正的未观测变量相关。计划评估文献涉及这样一个问题,即一个人永远不会在接受治疗和不接受治疗的情况下观察个人结果。在这些情况下,有必要确定假设,以克服因缺少将被称为主要数据集的信息而导致的识别不足。这种识别问题的一个常见解决方案是假设在条件独立的假设下可以使用辅助数据源恢复丢失的信息。识别策略的关键要素是辅助数据集必须提供关于在给定一组代理变量的情况下感兴趣的真实变量的条件分布的信息,其中在主样本和辅助样本中都观察到代理变量。该项目通过广义非线性矩模型方法得到了参数估计的半参数效率方差界,其中采样信息由主样本和辅助样本组成。在原始数据集中不能直接观察到矩条件中感兴趣的变量。主数据集包含与感兴趣的变量相关的代理变量。另一方面,辅助数据集包含关于给定代理变量的感兴趣变量的条件分布的信息。通过假设该条件分布在主数据集中和辅助数据集中都是相同的来实现识别。本项目所得到的结果既适用于两个样本独立的“检验出样本”的情况,也适用于“样本中检验”的情况,其中辅助样本是主样本的子集。基于Sieve的半参数估计器被用来在倾向性得分未知、倾向性已知或假定倾向性属于正确指定的参数族的情况下达到半参数效率界。这些估计量只使用条件期望的一个非参数估计,而不需要矩函数的条件期望和倾向得分的两个非参数估计。它们需要比文献中现有的更弱的正则性条件。它们还允许条件变量和非光滑矩条件的无界支持,并且不要求倾向性得分函数必须从零和一一致有界,这些结果将被推广到因变量或条件变量是带误差测量的条件矩模型。在这些情况下,在辅助数据集中只能观察到被怀疑有误差地测量的变量的子集。目前可用的估计量需要知道半参数效率方差界。第二个扩展将考虑基于非参数最大似然原理的估计器,这些估计器在不知道其特定形式的情况下达到半参数效率界。将进行广泛的蒙特卡罗模拟和实证说明,以评估竞争估计器的有限样本效率影响。拟议的项目涉及与纽约大学的陈晓红教授和杜克大学的亚历山德罗·塔罗齐教授的联合工作。广泛的影响:该项目开发的结果将适用于各种模型,包括非经典测量误差模型、缺失数据模型和非线性治疗效应模型。这项提议是计量经济学行业更大的研究议程的一部分,目的是开发方法来估计具有潜在变量的模型。经济学中的许多经验研究因相关变量的存在而变得复杂,这些变量没有被观察到,通常是因为它们只能以不完整或腐败的方式获得。重要的例子包括面板数据分析中的损耗和测量误差的存在,这些误差可能与真实的未观测变量相关。另一个例子是项目评估文献,其中对治疗效果的估计必须克服这样一个事实,即一个人永远不会观察有没有治疗的个人结果。在这种情况下,基于条件独立关系的识别假设变得必要,以克服由于原始数据集中的缺失信息而导致的识别不足。拟议的项目还将为调查数据集的设计提供有用的指导,这些数据集为经济计量模型的分析提供关键的数据输入。
英文摘要
Many empirical studies in economics are complicated by the presence of relevant variables that are not observed, either because they are only available in an incomplete or corrupted way, or because they are unobservable by their own nature. Important examples include attrition in panel data analysis and the ubiquitous presence of measurement error which can potentially be correlated with the true unobserved variables. The program evaluation literature is concerned with the issue that one never observes individual outcomes with and without treatment. In these circumstances, identifying assumptions become necessary to overcome the lack of identification that results from the missing information in what will be referred to as the primary data set. One common solution to this identification problem is the assumption that the missing information can be recovered using auxiliary data sources under a conditional independence assumption. The key element of the identification strategy is that the auxiliary data set must provide information about the conditional distribution of the true variables of interest given a set of proxy variables, where the proxy variables are observed in both the primary sample and the auxiliary sample. This project derives semiparametric efficiency variance bounds for the estimation of parameters defined through generalized nonlinear method of moment models, where the sampling information consists of a primary sample and an auxiliary sample. The variables of interest in the moment conditions are not directly observable in the primary data set. The primary data set contains proxy variables which are correlated with the variables of interest. On the other hand, the auxiliary data set contains information about the conditional distribution of the variables of interest given the proxy variables. Identification is achieved by the assumption that this conditional distribution is the same in both the primary and auxiliary data sets.The results derived in this project are applicable to both the "verify-out-of-sample" case, where the two samples are independent, and the "verify-in-sample case", where the auxiliary sample is a subset of the primary sample. Sieve based semiparametric estimators are developed to achieve the semiparametric efficiency bounds when the propensity score is unknown, when the propensity is known, or when the propensity is assumed to belong to a correctly specified parametric family. These estimators only use one nonparametric estimate of conditional expectation and do not require two nonparametric estimates of both the conditional expectation of the moment functions and the propensity score. They require weaker regularity conditions than the existing ones in the literature. They also allow for unbounded support of conditional variables and nonsmooth moment conditions, and do not require the strong assumption that the propensity score function has to be uniformly bounded away from zero and one.These results will be extended to conditional moment models in which either the dependent variables or the conditioning variables are measured with errors. In these cases only a subset of the variables that are suspected to be measured with error are observable in the auxiliary data set. The estimators currently available require knowledge of the semiparametric efficiency variance bounds. A second extension will consider estimators based on nonparametric maximum likelihood principles that achieve the semiparametric efficiency bound without knowledge of its particular form. Extensive monte carlo simulations and an empirical illustration will be performed to evaluate the finite sample efficiency implications of competing estimators. The proposed project involves joint work with Professor Xiaohong Chen from New York University and Professor Alessandro Tarozzi from Duke University.Broader Impact: The results developed in this project will be applicable to a wide variety of models, including non classical measurement error models, missing data models and nonlinear treatment effect models. This proposal is part of a larger research agenda in the econometrics profession to develop methods to estimate models with latent variables. Many empirical studies in economics are complicated by the presence of relevant variables that are not observed, usually because they are only available in incomplete or corrupted ways. Important examples include attrition in panel data analysis and the presence of measurement error which can potentially be correlated with the true unobserved variables. Another example is the program evaluation literature, where the estimation of treatment effects has to overcome the fact that one never observes individual outcomes with and without treatment. In such circumstances, identifying assumptions based on conditional independence relations become necessary to overcome the lack of identification that results from the missing information in the primary data set. The proposed project will also provide useful guidance to the design of survey data sets, which generate the crucial data input for the analysis of econometric models.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Numerical Bootstrap and Constrained Estimation
  • 批准号:
    1658950
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.43万
  • 财政年份:
    2017
  • 负责人:
    Han Hong
  • 依托单位:
A Computational Implementation of GMM
  • 批准号:
    1459975
  • 项目类别:
    Standard Grant
  • 资助金额:
    $18.3万
  • 财政年份:
    2015
  • 负责人:
    Han Hong
  • 依托单位:
Efficient Resampling and Simulation Methods for Nonlinear Econometric Models
  • 批准号:
    1325805
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.67万
  • 财政年份:
    2013
  • 负责人:
    Han Hong
  • 依托单位:
Collaborative Research: Statistical Properties of Numerical Derivatives and Algorithms
  • 批准号:
    1024504
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.73万
  • 财政年份:
    2010
  • 负责人:
    Han Hong
  • 依托单位:
海外基金