课题基金 / 基金详情

A Unifying Framework for High-Dimensional Additive Modeling

A Unifying Framework for High-Dimensional Additive Modeling
高维加性建模的统一框架
批准号:
1915855
负责人:
Noah Simon
金额:
$22.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-01 至 2023-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
最近的技术进步为研究人员提供了丰富的数据。要释放这些庞大数据源的潜力,通常需要将一个反应变量(例如,疾病的存在或对治疗的积极反应)与大量潜在的感兴趣特征(例如,基因表达水平或大脑区域的活动)联系起来。在少数受试者(或患者)上测量大量特征(数千或更多)越来越普遍。通常,这些特征中只有一小部分与响应变量相关。为了确定这些相关特征,使用统计/机器学习建模算法来自动选择最能预测响应的一小部分特征。这些算法中的许多都对它们所允许的模型施加了严格的限制;如线性。在科学领域,只有少量的受试者被测量,这些限制可能是有用的:构建更复杂的模型需要更多的受试者。然而,它们有时过于严格。在这个项目中,将开发一个框架来估计在高维问题中采用变量选择的限制性较小的(加性)模型。这个框架将有助于克服许多计算方面的挑战。它还将为分析这种高维估计器的统计行为(包括有限计算资源的影响)奠定理论基础。还将开发用于灵活高维建模的公开可用的软件实现。这个项目涉及非参数估计和惩罚回归的开创性问题。一般来说,高维稀疏非参数回归的计算和理论挑战是分开研究的:构建迭代算法,最终达到预定的最小公差范围内,同时研究精确最小器的统计性质(如收敛速度)。此外,对于非参数问题,现有的理论研究往往集中在目标隐含的结构(如稀疏性)完全成立时的统计性质上。本项目旨在融合计算和统计最优性的研究,在高维可加性和更一般的非参数模型的设置。更具体地说,它的目的是分析近似的统计性质,而不是精确的,最小化;从那里,它描述了获得具有最优性保证的估计器所需的下降迭代次数。此外,该项目旨在将这些想法扩展到结构/平滑可能被错误指定的设置中。为了应对这些挑战,该项目汇集了来自凸优化、经验过程理论、惩罚回归和近似理论的思想,并将作为将这些知识体系结合在一起的模板。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Recent technological advances have provided researchers with a wealth of data. Unlocking the potential of these vast data sources often involves linking a response variable (e.g., the presence of a disease, or a positive response to a treatment) to a large number of potential features of interest (e.g. gene expression levels or activities of brain regions). It is increasingly common to measure a large number of features (thousands or more), on a small number of subjects (or patients). Generally, only a small number of these features are related to the response variable. To determine those relevant features, statistical/machine-learning modelling algorithms are used to automatically select a small subset of features that are most predictive of response. Many of these algorithms enforce strong restrictions on the models they allow; e.g. linearity. In scientific domains, where only a small number of subjects are measured, these restrictions can be useful: building more complex models requires more subjects. However, they are sometimes overly restrictive. In this project, a framework to estimate less restrictive (additive) models that employ variable selection in high-dimensional problems will be developed. This framework will help overcome a number of computational challenges. It will additionally lay a theoretical foundation for analyzing the statistical behavior of such high dimensional estimators (that will include the impact of finite computational resources). A publicly available software implementation for flexible high-dimensional modeling will also be developed.This project engages seminal questions in nonparametric estimation and penalized regression. Generally, the computational and theoretical challenges of sparse nonparametric regression in high dimensions are studied separately: Iterative algorithms are constructed that eventually get within a prespecified tolerance of the minimum, while statistical properties (e.g. convergence rates) of the exact minimizer are studied. In addition, for non-parametric problems, existing theoretical studies have often focused on statistical properties when the structure implied by the objective (e.g., sparsity) holds exactly. This project aims to merge the study of computational and statistical optimality, in the setting of high-dimensional additive, and more general nonparametric, models. More specifically, it aims to analyze the statistical properties of approximate, rather than exact, minimizers; from there, it characterizes the number of descent iterations needed to obtain estimators with optimality guarantees. In addition, the project aims to extend these ideas to settings where the structure/smoothness may be misspecified. To address these challenges, the project brings together ideas from convex optimization, empirical process theory, penalized regression, and approximation theory, and will serve as a template for engaging those bodies of knowledge together.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2021-05
期刊: J. Mach. Learn. Res.
影响因子: --
作者: [Yunhua Xiang;Tianyu Zhang;Xu Wang;A. Shojaie;N. Simon]
通讯作者: Yunhua Xiang;Tianyu Zhang;Xu Wang;A. Shojaie;N. Simon
DOI: --
发表时间: 2019-03
期刊: Journal of machine learning research : JMLR
影响因子: --
作者: [Asad Haris;N. Simon;A. Shojaie]
通讯作者: Asad Haris;N. Simon;A. Shojaie
DOI: 10.1214/23-ejs2188
发表时间: 2022-06
期刊: Electronic Journal of Statistics
影响因子: 1.1
作者: [Tianyu Zhang;N. Simon]
通讯作者: Tianyu Zhang;N. Simon
海外基金