课题基金 / 基金详情

RUI: Predictive models with Incomplete and Fragmented Observations, and New Advances in Virtual Re-sampling for Big Data

RUI: Predictive models with Incomplete and Fragmented Observations, and New Advances in Virtual Re-sampling for Big Data
RUI:具有不完整和碎片观测的预测模型,以及大数据虚拟重采样的新进展
批准号:
2310504
负责人:
Majid Mojirsheibani
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31

项目摘要

项目成果

Majid Mojirsheibani的其他基金

相似基金

相关文献

中文摘要
翻译
该项目的一个主要重点是开发新的程序,以便在缺少数据的情况下进行统计建模、预测和推理。在医学、工程学、经济学和社会科学的许多领域,普遍存在不完整、缺失、删减和部分观察到的数据,这反过来又使数据驱动的决策过程中的预测和推理任务复杂化。研究者将研究和探索几种处理复杂数据结构中缺失值的新方法的有效性,而不会对导致信息缺失的潜在机制施加不切实际或不必要的严格条件。本研究项目的另一个主要目标是开发有效的数据重新采样方法,以减轻大数据场景中计算机密集型统计方法的巨大计算成本,在大数据场景中,数据分析师必须处理和整理大量数据。这种高效方法的出现是及时的,因为超大数据集的浪潮已经接管了医学、农业和环境保护领域的许多数据分析计划。此外,该项目还为研究生和本科生提供研究经验,其中许多人将被说服继续在STEM学科进行进一步的学习和研究。本研究项目涉及与预测模型和推理相关的两大类问题。第一部分侧重于预测模型中的选定主题,如回归和分类,用于许多非标准的现实设置。具体来说,研究人员将在一般度量空间中为不完整和分数观察数据开发几个局部平均型回归估计器,并将其应用于统计分类和无监督机器学习的相关问题。目的是对这些估计量在各种范数下的收敛性进行严格的研究,这是正确预测和推理所必需的。特别是,本项目将研究和开发新的指数性能界的估计器的Lp范数。研究了不完整和碎片化功能数据的带宽估计问题;这一点尤其重要,因为像MISE或ISE这样的最佳带宽最小化量在分类中不一定是最佳的。本研究计划的第二部分考虑了虚拟重采样的新目标,作为一种方法来减少大数据自举在许多重要和具有挑战性的问题上的巨大计算成本,同时仍然保留了自举方法的优点。特别是,研究者将开发虚拟重采样策略,以(i)近似大数据场景中多个测试问题的几个精炼的高批评统计数据的分布;(ii)在双样本问题中加速密度和回归估计器的重要泛函的对数缓慢的收敛速度,例如在大数据场景中基于反卷积密度估计器及其误差变量模型的超泛函的问题。为了实现(i)和(ii)项下的目标,研究者将使用文献中bootstrap经验过程的强近似中使用的方法。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
A major focus of this project is on the development of new procedures to carry out statistical modeling, prediction, and inference in the presence of missing data. Incomplete, missing, censored, and partially observed data are prevalent in many areas of medical sciences, engineering, economics and social sciences, which can in turn complicate the task of prediction and inference in data-driven decision-making processes. The investigator will study and explore the effectiveness of several new methods for handling missing values in complex data structures without imposing unrealistic or unnecessarily stringent conditions on the underlying mechanisms that cause the absence of information. Another major aim of this research project is to develop efficient data re-sampling methods to alleviate the formidable computational cost of computer-intensive statistical methods in big-data scenarios, where the data analyst must deal with, and sort through, massive amounts of data. The advent of such efficient methods is timely as the wave of ultra-large datasets has taken over many data-analytic initiatives in medicine, agriculture, and environmental protection. Additionally, this project embraces research experiences for graduate and undergraduate students, many of whom will then be persuaded to move on to further studies and research careers in STEM disciplines.This research project deals with two broad classes of problems related to predictive models and inference. The first part focuses on selected topics in predictive models such as regression and classification for a number of nonstandard realistic setups. Specifically, the investigator will develop several local-averaging-type regression estimators in general metric spaces for incomplete and fractionally observed data with applications to statistical classification and the related problem of unsupervised machine learning. The aim is to carry out a rigorous study of the convergence properties of these estimators in various norms which is necessary for correct prediction and inference. In particular, this project will study and develop new exponential performance bounds for the Lp norms of the proposed estimators. The problem of bandwidth estimation for incomplete and fragmented functional data will also be studied; this is particularly important as the optimal bandwidth minimizing quantities such as the MISE or ISE is not necessarily optimal in classification. The second part of this research plan considers new objectives in virtual re-sampling as a method to reduce the formidable computational cost of big-data bootstrap in a number of important and challenging problems, while still retaining the benefits of bootstrap methodology. In particular, the investigator will develop virtual re-sampling strategies to (i) approximate the distribution of several refined higher criticism statistics for multiple testing problems in big-data scenarios, and (ii) to speed up the logarithmically slow rates of convergence of important functionals of density and regression estimators in two-sample problems such as those based on deconvolution density estimators and their sup-functionals for errors-in-variables models in big-data scenarios. To achieve the objectives under (i) and (ii), the investigator will use adaptations of the methodologies used in the strong approximations of bootstrap empirical processes in the literature.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RUI: Partially Observed Curves, and Big-Data Virtual Bootstrap
RUI: Classification, regression, and density estimation with missing variables
海外基金