RUI: Classification, regression, and density estimation with missing variables
RUI: Classification, regression, and density estimation with missing variables
批准号:
1407400
负责人:
Majid Mojirsheibani
金额:
$12.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2018-08-31
中文摘要
该项目发展了在存在缺失或不完整数据的情况下用于非参数分类和曲线估计的统计理论和方法。许多数据集缺少值;这些数据包括来自生物医学研究、遥感以及社会科学的数据。有许多处理丢失数据的经典方法。许多现有的结果首先归因于缺失值,然后应用标准的统计技术进行推断。然而,由于失去了数据中的独立性假设,对这类技术的理论有效性的研究可能会变得困难;对于无分布的统计方法来说尤其如此。首席研究员(PI)的新结果将回答统计分类和模式识别中的一些基本问题,并应用于生物医学、遥感和社会科学。新的结果还将解决机器学习和统计分类交叉的许多重要理论问题。缺失协变量分类中的一个长期存在的问题涉及到数据和新的未分类观测中都可能出现缺失协变量的情况。这从根本上不同于更简单的问题,即丢失的协变量只出现在数据中。在后一种情况下,可以使用基于Horvitz-Thompson逆加权的标准方法来构造渐近最优分类器。PI研究项目的一部分集中在这一具有挑战性的缺失协变量的分类案例上。PI将开发新的渐近最优的局部平均类型的分类器,如核规则和划分规则。这个项目的另一部分集中在PI先前在组合分类和估计方面的努力的延续和完善,基于最近在文献中获得的结果。PI目前正在开发新的方法来组合几个单独的分类器,使得得到的分类器的渐近误差将至少与最好的单独分类器的渐近误差一样好。PI还将开发以最佳方式组合几个回归函数估计器的方法。将使用经验过程理论中的工具来确定所得到的分类器和估计器的大样本最优性。第三部分研究了缺失数据下核密度估计的各种范数的弱收敛问题。PI将研究这些统计量的加权自举近似。这样的结果将允许某人在存在缺失值的情况下为未知密度构造正确的置信带。这里的主要工具是强逼近定理,它允许人们用一系列布朗桥来代替加权自举经验过程。
英文摘要
This project develops statistical theory and methods for nonparametric classification and curve estimation in the presence of missing or incomplete data. Many data sets have missing values; these include the data from biomedical studies, remote sensing, as well as social sciences. There are a number of classical approaches for handing the missing data. Many of the existing results first impute for the missing values and then apply a standard statistical technique to carry out inferences. However, a study of the theoretical validity of such techniques can become intractable due to the loss of independence assumption in the data; this is particularly true for distribution-free statistical methods. The new results of the Principal Investigator (PI) will answer a number of fundamental questions in statistical classification and pattern recognition with applications to biomedical, remote sensing, and social sciences. The new results will also solve many important theoretical problems at the intersection of machine learning and statistical classification.A long-standing problem in classification with missing covariates involves the situation where missing covariates can appear in both the data and in the new unclassified observation. This is fundamentally different from the simpler problem where missing covariates appear in the data only. In the latter case, standard methods based on Horvitz-Thompson inverse weighting can be used to construct asymptotically optimal classifiers. One part of the PI's research project focuses on this challenging case of classification with missing covariates. The PI will develop new asymptotically optimal local-averaging-type classifiers, such as kernel and partitioning rules. Another part of this project concentrates on the continuation and refinements of the PI's previous efforts on combined classification and estimation, based on recently obtained results in the literature. The PI is currently developing new methods to combine several individual classifiers in such a way that the asymptotic error of the resulting classifier will be at least as good as that of the best individual classifier. The PI will also develop methods to combine several regression function estimators in an optimal way. Tools from the empirical process theory will be used to establish the large-sample optimality of the resulting classifiers and estimators. The third part of the project focuses on the weak convergence of various norms of kernel density estimates in the presence of missing data. The PI will study weighted bootstrap approximations of these statistics. Such results will allow someone to construct correct confidence bands for the unknown density in the presence of missing values. The main tools here are the strong approximation theorems that allow one to replace the weighted bootstrapped empirical processes by a sequence of Brownian bridges.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RUI: Predictive models with Incomplete and Fragmented Observations, and New Advances in Virtual Re-sampling for Big Data
-
批准号:2310504
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2023
-
负责人:Majid Mojirsheibani
-
依托单位:
RUI: Partially Observed Curves, and Big-Data Virtual Bootstrap
-
批准号:1916161
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2019
-
负责人:Majid Mojirsheibani
-
依托单位:
海外基金