课题基金 / 基金详情

Hypothesis Testing in High Dimensions Without Sparsity

Hypothesis Testing in High Dimensions Without Sparsity
无稀疏性的高维假设检验
批准号:
1712481
负责人:
Jelena Bradic
金额:
$12.5万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2022-08-31

项目摘要

项目成果

Jelena Bradic的其他基金

相似基金

相关文献

中文摘要
翻译
多样化科学数据的指数级增长代表着在复杂科学和工程领域取得实质性进展的前所未有的机会,例如发现新材料或药物。大规模收集和分析信息对社会有明显的好处:它可以帮助企业优化在线商务,医务工作者解决公共卫生问题,政府中断恐怖活动。该项目旨在设计严格的统计框架和先进的工具,用于传播和管理复杂物理和工程系统建模和设计中的不确定性。该项目的最终目标是通过将系统中的科学领域的数据与极高维空间中的多尺度和不确定参数相关联来促进重要的假设生成和加速发现,包括例如航空航天,工程,神经科学,基因-蛋白质疾病网络,材料科学,气候科学,自治系统和网络物理系统或自组织生物系统。该项目将使系统的设计具有可验证的属性;揭示如何在各种应用中评估结果的可信度;允许对结论的不确定性和实际实验和验证的敏感性进行探索。了解我们周围复杂且日益数据密集的世界依赖于构建强大的经验模型,即,真实的复杂系统的表示,使决策者能够预测行为并回答“假设”问题。 新的统计研究是必要的,以处理潜在的高维空间的不确定参数,积极的多物理耦合,以及模型本身的不确定性。 此外,对于这些大系统,不确定性条件下的决策还没有基本的理论。为了满足这些需求,该项目打算开发以下功能:用于逆建模的新方法,以扩展到高维多尺度/多物理系统;对物理模型本身的不确定性和不足之处的可量化和可概括的理解,以及对模型错误和假设非常鲁棒的全新决策范式。特别是,新的决策理论的高维回归模型将被开发,是高度和可证明的鲁棒性的模型误指定,包括丢失,非稀疏和大规模(爆炸)的数据结构。 因此,这些方法将是理解密集高维模型的基本限制的第一次尝试。此外,将在无法进行完美估计或变量选择的情况下研究删失分析模型(对数秩、加性风险、比例风险和竞争风险);这些研究将能够弥合高维情况下实践与理论之间的差距。此外,还将开发一个新的框架,用于汇总可能相互关联和随时间变化的异质和不同的数据结构。在分析从一个固定的观测源到许多移动的观测源的不同数据时,必须小心地用简约来平衡增加的灵活性。
英文摘要
The exponential growth of diverse scientific data represents an unprecedented opportunity to make substantial advances in complex science and engineering, such as the discovery of novel materials or drugs. The collection and analysis of information on massive scales has clear benefits for society: it can help businesses optimize online commerce, medical workers address public health issues, and governments interrupt terrorist activities. This project aims at designing rigorous statistical framework and advanced tools for propagating and managing uncertainty in the modeling and design of complex physical and engineering systems. The ultimate goal of this project is to facilitate significant hypothesis generation and accelerate discovery by correlating data across scientific domains in systems with multi-scale and uncertain parameters in extremely high-dimensional spaces, including for example aerospace, engineering, neuroscience, gene-protein disease networks, materials science, climate science, autonomous systems, and cyber-physical systems or self-organized biological systems. This project will enable the design of systems with verifiable properties; reveal how to value the trustworthiness of results in a wide variety of applications; allow conclusions to be probed for their sensitivity to uncertainties and practical experimentation and validation. Understanding the complex and increasingly data-intensive world around us relies on the construction of robust empirical models, i.e., representations of real, complex systems that enable decision makers to predict behaviors and answer "what-if" questions. Novel statistical research is needed for dealing with the underlying high dimensionality of the space of uncertain parameters, active multi-physics coupling, and uncertainty in the models themselves. Besides, there is no fundamental theory for decision making under uncertainty for these large-scale systems. To address these needs, this project intends to develop the following capabilities: new methods for inverse modeling to scale to high-dimensional multi-scale/multi-physics systems; a quantifiable and generalizable understanding of uncertainties and inadequacies in the physical models themselves and entirely new paradigms for decision making that is extremely robust to model misspecification and assumptions. In particular, new decision theory for high-dimensional regression models will be developed that is highly and provably robust to the model misspecification, including missing, non-sparse and large-scale (exploding) data structures. As such, these methods will be the first attempts at understanding the fundamental limitations of dense high-dimensional models. Moreover, censored analysis models (log-rank, additive hazard, proportional hazard and competing risk) will be studied under the setting where perfect estimation or variable selection is not possible; these studies will be able to bridge the gap between the practice and theory in the high-dimensional setting. Furthermore, a new framework for aggregating heterogeneous and disparate data structures that may be correlated and time-dependent will be developed. The added flexibility in analyzing disparate data going from one stationary source to many moving sources of observations must be carefully balanced with parsimony.
期刊论文(13)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1080/01621459.2017.1356319
发表时间: 2016-10
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Yinchu Zhu;Jelena Bradic]
通讯作者: Yinchu Zhu;Jelena Bradic
DOI: 10.1080/01621459.2020.1840989
发表时间: 2020-10
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Lan Wang;Bo Peng;Jelena Bradic;Runze Li;Y. Wu]
通讯作者: Lan Wang;Bo Peng;Jelena Bradic;Runze Li;Y. Wu
Censored Quantile Regression Forest
截尾分位数回归森林
DOI: --
发表时间: 2020
期刊: Proceedings of Machine Learning Research
影响因子: --
作者: [Li, Alexander Hanbo, Bradic, Jelena]
通讯作者: Bradic, Jelena
DOI: 10.1080/01621459.2016.1273116
发表时间: 2015-10
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Alexander Hanbo Li;Jelena Bradic]
通讯作者: Alexander Hanbo Li;Jelena Bradic
共 12 条
    Regularization for High Dimensional Inference and Sparse Recovery
    • 批准号:
      1205296
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $12.0万
    • 财政年份:
      2012
    • 负责人:
      Jelena Bradic
    • 依托单位:
    海外基金