课题基金 / 基金详情

RUI: Partially Observed Curves, and Big-Data Virtual Bootstrap

RUI: Partially Observed Curves, and Big-Data Virtual Bootstrap
RUI:部分观察曲线和大数据虚拟引导程序
批准号:
1916161
负责人:
Majid Mojirsheibani
金额:
$17.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-01 至 2023-08-31

项目摘要

项目成果

Majid Mojirsheibani的其他基金

相似基金

相关文献

中文摘要
翻译
科学学科中的许多真实数据集,如生物医学、工程和社会科学,包含缺失、删节或部分观测值,这可能使统计估计和推理的任务变得更加复杂。该研究项目的一部分侧重于开发新的灵活的统计方法,以便在数据不完整和缺失的情况下进行准确的预测和推断。在这里,数据可以是高维的,也可以是功能性的,其中每个数据值可以是一条曲线。在这个研究项目的另一部分,PI考虑开发新的高效的计算机密集型方法来处理大数据场景,其中数据规模可能太大而无法调用经典方法。近年来,大数据已成为当前的研究前沿之一,学术界和工业界对大数据驱动的决策程序越来越感兴趣。在这一领域仍然存在许多计算和理论挑战,需要新的方法。PI的新方法将解决机器学习和统计推断交叉的许多重要统计问题。该研究涉及三大类与非标准设置中的预测和推理相关的问题。其中包括当协变量曲线在其域的某些子集上可能不可观察时的功能分类问题。然而,与文献中的一些早期结果不同,PI的方法并没有对导致信息缺失或审查的机制施加任何随机缺失(MAR)类型的假设。该方法允许不完全协变量出现在新的未分类曲线和数据中。给定观察到的协变量片段,目的是基于局部平均方法构建强一致性非参数分类器。第二类问题处理响应变量缺失情况下核回归估计量的一致渐近性问题。这被普遍认为是一个难题。这些估计量的最大偏差的极限分布可用于构造渐近正确的均匀置信带,或用于对未知回归函数进行拟合优度检验。在这里,PI将考虑MAR和不可忽略的缺失响应假设。第三组问题侧重于大数据场景下新的加权自举方法的开发。PI的方法旨在减少与大数据重复采样相关的计算负担,同时仍然保留了bootstrap方法的优点。所开发的方法将用于在大数据场景中更好地近似核和反卷积密度估计器的抽样分布,以及它们的重要函数(如sup- norm和lp -norm)。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Many real data sets in scientific disciplines, such as biomedical, engineering, and social sciences, contain missing, censored, or partially observed values, and this can make the task of statistical estimation and inference significantly more complicated. Part of this research project focuses on the development of new flexible statistical methods to perform accurate prediction and inference in the presence of incomplete and missing data. Here, the data could be high-dimensional as well as functional, where each data value can be a curve. In another part of this research project, the PI considers the development of new efficient computer-intensive methods to deal with Big-data scenarios, where the data size may be too large to invoke classical approaches. Big-data has been one of the current research frontiers in recent years and there has been a growing interest in Big-data-driven decision-making procedures in both academia and the industry. There are still many computational and theoretical challenges in this area that require new methodologies. The PI's new approaches will solve a number of important statistical problems at the intersection of machine learning and statistical inference.The research deals with three broad classes of problems related to prediction and inference in some nonstandard setups. These include the problem of functional classification when the covariate curves may be unobservable on some subsets of their domain. However, unlike some of the earlier results in the literature, the PI's approach does not impose any missing-at-random (MAR) type assumptions on the mechanisms that cause the absence or censoring of information. The approach allows for incomplete covariate to appear in the new unclassified curves as well as in the data. Given the observed covariate fragments, the aim is to construct strongly consistent nonparametric classifiers based on local-averaging methods. The second class of problems deals with uniform asymptotics for kernel regression estimators in the presence of missing response variables. This is generally acknowledged to be a difficult problem. The limiting distribution of the maximal deviation of such estimators can be used to construct asymptotically correct uniform confidence bands, or to perform goodness-of-fit tests, for an unknown regression function. Here, the PI will consider both MAR and non-ignorable missing response assumptions. The third set of problems focuses on the development of new weighted bootstrap methods for Big-data scenarios. The PI's approach aims at reducing the computational burden associated with the repeated sampling of Big-data, while still retaining the benefits of bootstrap methodology. The developed methods will be used to better approximate the sampling distribution of kernel and deconvolution density estimators, as well as their important functionals (such as sup- and Lp-norms), in the Big-data scenario.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1007/s00184-020-00794-y
发表时间: 2020-08-19
期刊: METRIKA
影响因子: 0.7
作者: [Han,Eric, Mojirsheibani,Majid]
通讯作者: Mojirsheibani,Majid
A nearest-neighbor-based ensemble classifier and its large-sample optimality
基于最近邻的集成分类器及其大样本最优性
DOI: 10.1080/00949655.2021.1882458
发表时间: 2021
期刊: Journal of Statistical Computation and Simulation
影响因子: 1.2
作者: [Mojirsheibani, Majid, Pouliot, William]
通讯作者: Pouliot, William
DOI: 10.1016/j.jmva.2021.104755
发表时间: 2021-03
期刊: J. Multivar. Anal.
影响因子: --
作者: [M. Mojirsheibani]
通讯作者: M. Mojirsheibani
DOI: 10.1007/s00184-023-00923-3
发表时间: 2022-12
期刊: Metrika
影响因子: 0.7
作者: [M. Mojirsheibani;William Pouliot;Andre Shakhbandaryan]
通讯作者: M. Mojirsheibani;William Pouliot;Andre Shakhbandaryan
共 7 条
    RUI: Predictive models with Incomplete and Fragmented Observations, and New Advances in Virtual Re-sampling for Big Data
    RUI: Classification, regression, and density estimation with missing variables
    国内基金
    海外基金
    基于分数阶衍射的PT及Partially-PT对称非线性系统中的空间孤子研究
    • 批准号:
      11764022
    • 项目类别:
      地区科学基金项目
    • 资助金额:
      33.0万元
    • 批准年份:
      2017
    • 负责人:
      黎磊
    • 依托单位: