RUI: Classification, regression, and density estimation with missing variables
RUI: Classification, regression, and density estimation with missing variables
批准号:
1407400
负责人:
Majid Mojirsheibani
金额:
$12.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2018-08-31
中文摘要
本项目主要研究在数据缺失或不完整的情况下,非参数分类和曲线估计的统计理论和方法。许多数据集都有缺失值,其中包括生物医学研究、遥感和社会科学的数据。有许多经典的方法来处理缺失数据。 许多现有的结果首先归咎于缺失值,然后应用标准的统计技术进行推断。 然而,这种技术的理论有效性的研究可能会变得棘手,由于在数据中的独立性假设的损失,这是特别真实的分布自由的统计方法。主要研究者(PI)的新成果将回答统计分类和模式识别中的许多基本问题,并应用于生物医学、遥感和社会科学。 新的结果也将解决机器学习和统计分类交叉点上的许多重要理论问题。在具有缺失协变量的分类中,一个长期存在的问题涉及缺失协变量可能同时出现在数据和新的未分类观察中的情况。这与更简单的问题有根本的不同,在更简单的问题中,缺失的协变量只出现在数据中。在后一种情况下,基于Horvitz-Thompson逆加权的标准方法可用于构造渐近最优分类器。PI的研究项目的一部分集中在这个具有挑战性的分类案例中,缺少协变量。PI将开发新的渐近最优局部平均型分类器,如内核和分区规则。 本项目的另一部分集中在PI的结合分类和估计,最近获得的结果在文献的基础上,以前的努力的延续和改进。 PI目前正在开发新的方法,以联合收割机结合几个单独的分类器,以这种方式,所产生的分类器的渐近误差将至少是最好的个人分类器。 PI还将开发以最佳方式联合收割机组合几个回归函数估计器的方法。从经验过程理论的工具将被用来建立大样本最优的分类器和估计。该项目的第三部分重点讨论了在缺失数据的情况下,核密度估计的各种范数的弱收敛性。PI将研究这些统计量的加权bootstrap近似。这样的结果将允许某人在存在缺失值的情况下为未知密度构建正确的置信带。这里的主要工具是强逼近定理,它允许用一系列布朗桥来代替加权自举经验过程。
英文摘要
This project develops statistical theory and methods for nonparametric classification and curve estimation in the presence of missing or incomplete data. Many data sets have missing values; these include the data from biomedical studies, remote sensing, as well as social sciences. There are a number of classical approaches for handing the missing data. Many of the existing results first impute for the missing values and then apply a standard statistical technique to carry out inferences. However, a study of the theoretical validity of such techniques can become intractable due to the loss of independence assumption in the data; this is particularly true for distribution-free statistical methods. The new results of the Principal Investigator (PI) will answer a number of fundamental questions in statistical classification and pattern recognition with applications to biomedical, remote sensing, and social sciences. The new results will also solve many important theoretical problems at the intersection of machine learning and statistical classification.A long-standing problem in classification with missing covariates involves the situation where missing covariates can appear in both the data and in the new unclassified observation. This is fundamentally different from the simpler problem where missing covariates appear in the data only. In the latter case, standard methods based on Horvitz-Thompson inverse weighting can be used to construct asymptotically optimal classifiers. One part of the PI's research project focuses on this challenging case of classification with missing covariates. The PI will develop new asymptotically optimal local-averaging-type classifiers, such as kernel and partitioning rules. Another part of this project concentrates on the continuation and refinements of the PI's previous efforts on combined classification and estimation, based on recently obtained results in the literature. The PI is currently developing new methods to combine several individual classifiers in such a way that the asymptotic error of the resulting classifier will be at least as good as that of the best individual classifier. The PI will also develop methods to combine several regression function estimators in an optimal way. Tools from the empirical process theory will be used to establish the large-sample optimality of the resulting classifiers and estimators. The third part of the project focuses on the weak convergence of various norms of kernel density estimates in the presence of missing data. The PI will study weighted bootstrap approximations of these statistics. Such results will allow someone to construct correct confidence bands for the unknown density in the presence of missing values. The main tools here are the strong approximation theorems that allow one to replace the weighted bootstrapped empirical processes by a sequence of Brownian bridges.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RUI: Predictive models with Incomplete and Fragmented Observations, and New Advances in Virtual Re-sampling for Big Data
-
批准号:2310504
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2023
-
负责人:Majid Mojirsheibani
-
依托单位:
RUI: Partially Observed Curves, and Big-Data Virtual Bootstrap
-
批准号:1916161
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2019
-
负责人:Majid Mojirsheibani
-
依托单位:
海外基金