课题基金 / 基金详情

FRG: Collaborative Research: Quantile-Based Modeling for Large-Scale Heterogeneous Data

FRG: Collaborative Research: Quantile-Based Modeling for Large-Scale Heterogeneous Data
FRG:协作研究:大规模异构数据的基于分位数的建模
批准号:
1952486
负责人:
Qi Zheng
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-06-01 至 2024-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
技术的快速发展导致了科学、经济、工程、医疗保健和许多其他学科中大规模异构数据的巨大增长。例如,在现代卫生信息系统中,电子健康记录通常收集来自不同疾病类别的异质人群的大量患者信息。这些数据为了解不同亚群的特征和结果之间的关系提供了独特的机会。现有的方法并没有完全解决计算和统计方面的巨大挑战。为了挖掘信息丰富的数据的真正潜力,该项目将为分析大规模异构数据提供新的计算和统计范式和坚实的理论基础。此外,该项目还将为研究生提供研究培训机会。该项目将建立一个统一的、基于分位数建模的框架,其总体目标是在分析异构数据时实现有效性和可靠性,特别是在潜在解释变量数量和样本量都很大的情况下。具体目标是:(1)为大规模异构数据开发基于重采样的推理;(2)开发贝叶斯算法和可扩展和可解释的结构感知方法,以更好地进行推理;(3)发展多协变量的分位数最优决策规则估计与推理;(4)提出了一种新的基于审查的大尺度分位数回归的估计和推理方法。该项目将解决数据大小和维度的可扩展性、异质性和结构的探索、对鲁棒性的需求以及利用不完整观测的能力方面的一些关键障碍。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The rapid development of technology has led to the tremendous growth of large-scale heterogeneous data in science, economics, engineering, healthcare, and many other disciplines. For example, in a modern health information system, electronic health records routinely collect a large amount of information on many patients from heterogeneous populations across different disease categories. Such data provide unique opportunities to understand the association between features and outcomes across different subpopulations. Existing approaches have not fully addressed the formidable computational and statistical challenges. To tap into the true potential of information-rich data, this project will develop a new computational and statistical paradigm and solid theoretical foundation for analyzing large-scale heterogeneous data. In addition the project will also provide research training opportunities for graduate students. The project will build a unified, quantile-modeling based framework with an overarching goal of achieving effectiveness and reliability in analyzing heterogeneous data, especially when both the number of potential explanatory variables and the sample size are large. The specific goals are (1) to develop resampling-based inference for large-scale heterogeneous data; (2) to develop Bayesian algorithms and scalable and interpretable structure-aware approach for better inference; (3) to develop quantile-optimal decision rule estimation and inference with many covariates; (4) to develop novel estimation and inference procedure for large-scale quantile regression under censoring. The project will address some of the key barriers in scalability to data size and dimensionality, exploration of heterogeneity and structures, need for robustness, and the ability to make use of incomplete observations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Globally Adaptive Longitudinal Quantile Regression with High Dimensional Compositional Covariates.
具有高维成分协变量的全局自适应纵向分位数回归
DOI: 10.5705/ss.202021.0006
发表时间: 2023-05
期刊: Statistica Sinica
影响因子: 1.4
作者: [Ma H, Zheng Q, Zhang Z, Lai H, Peng L]
通讯作者: Peng L
海外基金