课题基金 / 基金详情

RUI: Classification, regression, and density estimation with missing variables

RUI: Classification, regression, and density estimation with missing variables
RUI:分类、回归和缺失变量的密度估计
批准号:
1407400
负责人:
Majid Mojirsheibani
金额:
$12.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2018-08-31

项目摘要

项目成果

Majid Mojirsheibani的其他基金

相似基金

相关文献

中文摘要
翻译
本项目发展非参数分类和曲线估计在缺失或不完整的数据存在的统计理论和方法。许多数据集都有缺失值;这些数据包括来自生物医学研究、遥感以及社会科学的数据。有许多处理丢失数据的经典方法。许多现有的结果首先对缺失值进行估算,然后应用标准的统计技术进行推断。然而,由于数据中独立性假设的丧失,对这些技术的理论有效性的研究可能变得棘手;对于无分布的统计方法尤其如此。首席研究员(PI)的新结果将回答统计分类和模式识别方面的一些基本问题,并将其应用于生物医学、遥感和社会科学。新的结果也将解决机器学习和统计分类交叉领域的许多重要理论问题。缺失协变量分类中一个长期存在的问题是缺失协变量可能同时出现在数据和新的未分类观测中。这与缺少协变量只出现在数据中的简单问题根本不同。在后一种情况下,可以使用基于Horvitz-Thompson逆加权的标准方法来构造渐近最优分类器。PI研究项目的一部分集中在缺少协变量的分类这一具有挑战性的案例上。PI将开发新的渐近最优局部平均型分类器,如核规则和分区规则。这个项目的另一部分集中于PI先前在结合分类和估计方面的努力的延续和改进,基于最近在文献中获得的结果。PI目前正在开发新的方法来组合几个单独的分类器,使最终分类器的渐近误差至少与最佳单个分类器的渐近误差一样好。PI还将开发以最优方式组合几个回归函数估计器的方法。从经验过程理论的工具将被用来建立大样本最优的结果分类器和估计器。该项目的第三部分侧重于在缺失数据存在下核密度估计的各种规范的弱收敛性。PI将研究这些统计量的加权自举近似。这样的结果将允许某人在缺失值存在的情况下为未知密度构建正确的置信带。这里的主要工具是强逼近定理,它允许人们用一系列布朗桥来代替加权自举经验过程。
英文摘要
This project develops statistical theory and methods for nonparametric classification and curve estimation in the presence of missing or incomplete data. Many data sets have missing values; these include the data from biomedical studies, remote sensing, as well as social sciences. There are a number of classical approaches for handing the missing data. Many of the existing results first impute for the missing values and then apply a standard statistical technique to carry out inferences. However, a study of the theoretical validity of such techniques can become intractable due to the loss of independence assumption in the data; this is particularly true for distribution-free statistical methods. The new results of the Principal Investigator (PI) will answer a number of fundamental questions in statistical classification and pattern recognition with applications to biomedical, remote sensing, and social sciences. The new results will also solve many important theoretical problems at the intersection of machine learning and statistical classification.A long-standing problem in classification with missing covariates involves the situation where missing covariates can appear in both the data and in the new unclassified observation. This is fundamentally different from the simpler problem where missing covariates appear in the data only. In the latter case, standard methods based on Horvitz-Thompson inverse weighting can be used to construct asymptotically optimal classifiers. One part of the PI's research project focuses on this challenging case of classification with missing covariates. The PI will develop new asymptotically optimal local-averaging-type classifiers, such as kernel and partitioning rules. Another part of this project concentrates on the continuation and refinements of the PI's previous efforts on combined classification and estimation, based on recently obtained results in the literature. The PI is currently developing new methods to combine several individual classifiers in such a way that the asymptotic error of the resulting classifier will be at least as good as that of the best individual classifier. The PI will also develop methods to combine several regression function estimators in an optimal way. Tools from the empirical process theory will be used to establish the large-sample optimality of the resulting classifiers and estimators. The third part of the project focuses on the weak convergence of various norms of kernel density estimates in the presence of missing data. The PI will study weighted bootstrap approximations of these statistics. Such results will allow someone to construct correct confidence bands for the unknown density in the presence of missing values. The main tools here are the strong approximation theorems that allow one to replace the weighted bootstrapped empirical processes by a sequence of Brownian bridges.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RUI: Predictive models with Incomplete and Fragmented Observations, and New Advances in Virtual Re-sampling for Big Data
RUI: Partially Observed Curves, and Big-Data Virtual Bootstrap
海外基金