Collaborative Research: Use of Random Compression Matrices For Scalable Inference in High Dimensional Structured Regressions
Collaborative Research: Use of Random Compression Matrices For Scalable Inference in High Dimensional Structured Regressions
批准号:
2210206
负责人:
Aaron Scheffler
金额:
$11.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-06-15 至 2025-05-31
中文摘要
随着科学界进入数据驱动时代,利用大规模成像、基因和EHR数据来更好地描述和了解人类疾病,以改善治疗和预后,这是一个前所未有的机会。因此,使用灵活的统计模型分析这类数据集在过去十年中已成为一个非常活跃的研究领域。为此,该项目计划开发一类全新的方法,其基于的思想是对通过使用精心设计的机制压缩大数据而获得的数据集进行统计模型拟合。这一发展使得能够以前所未有的规模高效地对海量数据进行建模。虽然研究人员的动机主要来自对大量生物医学数据的复杂建模和不确定性量化,但统计方法足够普遍,足以在机器学习和环境科学的相关文献中留下重要的足迹。总体目标还包括开发软件工具包,以便更好地服务于相关学科的从业者。此外,这些项目将为研究生和本科生,包括女性和少数族裔社区的学生,提供最先进的统计方法和成像/遗传/电子健康记录数据的第一手培训机会。通过在高中生中以他们能理解的术语传播该项目的成果,该项目可以对提高公众的统计科学素养产生深远的影响。在复杂和高维数据时代,现代统计学习方法的两个关键方面是推理的准确性和规模。现代数据越来越复杂和高维,涉及的变量多,样本量大,不同变量之间的关系复杂。开发实用有效的(在存储和分析方面)和理论上“最优”的贝叶斯高维参数或非参数回归方法,以便从如此复杂的数据集中得出具有有效不确定性的准确推断是一个极其重要的问题。为了提供这个问题的一般解决方案,研究人员将开发基于使用少量随机线性变换的数据压缩的方法。该方法或者使用压缩来减少对应于每个变量的大量记录,在这种情况下,它保持特征解释以进行充分的推理,或者使用压缩来降低每个样本的协变量向量的维度,在这种情况下,重点仅放在响应的预测上。在任何一种情况下,在存在具有足够丰富的参数和非参数回归模型的高维数据的情况下,数据压缩促进了图形存储的高效、可伸缩和准确的贝叶斯推理/预测。一个重要的目标是建立关于压缩数据的拟合模型的收敛行为的精确的理论结果,该结果作为预测因子的数量、样本大小、随机线性变换的性质和这些模型的特征的函数。这些方法将结合大脑成像数据、基因数据和来自英国生物库数据库的电子健康记录(EHR)数据,用于研究神经疾病。该项目还将在更广泛的战线上为推进跨学科研究培训和扩大统计科学的参与做出贡献。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
As the scientific community moves into a data-driven era, there is an unprecedented opportunity to leverage large scale imaging, genetic and EHR data to better characterize and understand human disease to improve treatment and prognosis. Consequently, analysis of such datasets with flexible statistical models has become an enormously active area of research over the last decade. To this end, this project plans to develop a completely new class of methods, which are based on the idea of fitting statistical models on datasets obtained by compressing big data using a well designed mechanism. The development enables efficient modeling of massive data on an unprecedented scale. While the motivation of the investigators comes primarily from complex modeling and uncertainty quantification of massive biomedical data, the statistical methods are general enough to set important footprints in the related literature of machine learning and environmental sciences. The overarching goal also includes the development of software toolkits to better serve practitioners in related disciplines. Further, the projects will provide first hand training opportunities for graduate and undergraduate students, including female and students from minority communities, in state-of-the-art statistical methodologies and imaging/genetic/EHR data. By disseminating the outcome of the project among high school students in terminology that they can understand, the project can have far reaching effects to enhance public scientific literacy about statistics.Two crucial aspects of modern statistical learning approaches in the era of complex and high dimensional data are accuracy and scale in inference. Modern data are increasingly complex and high dimensional, involving a large number of variables and large sample size, with complex relationships between different variables. Developing practically efficient (in terms of storage and analysis) and theoretically “optimal” Bayesian high dimensional parametric or nonparametric regression methods to draw accurate inference with valid uncertainties from such complex datasets is an extremely important problem. To offer a general solution for this problem, the investigators will develop approaches based on data compression using a small number of random linear transformations. The approach either reduces a large number of records corresponding to each variable using compression, in which case it maintains feature interpretation for adequate inference, or, reduces the dimension of the covariate vector for each sample using compression, in which case the focus is only on prediction of the response. In either case, data compression facilitates drawing storage efficient, scalable and accurate Bayesian inference/prediction in presence of high dimensional data with sufficiently rich parametric and nonparametric regression models. An important goal is to establish precise theoretical results on the convergence behavior of the fitted models with compressed data as a function of the number of predictors, sample size, properties of random linear transformations and features of these models. The approaches will be used to study neurological disorders by combining brain imaging data, genetic data and electronic health records (EHR) data from the UK Biobank database. The project will also contribute on a broader front to advancing the interdisciplinary research training and broadening participation in statistical sciences.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: