课题基金 / 基金详情

CAREER: Calibrating Regularization for Enhanced Statistical Inference

CAREER: Calibrating Regularization for Enhanced Statistical Inference
职业:校准正则化以增强统计推断
批准号:
1753171
负责人:
Daniel McDonald
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-07-01 至 2020-06-30

项目摘要

项目成果

Daniel McDonald的其他基金

相似基金

相关文献

中文摘要
翻译
仅仅在十年前还难以想象的大量数据现在在科学研究中无处不在。投资决策基于数千种证券的价格,每微秒更新一次;大气科学家使用多分辨率卫星图像来了解气候变化;互联网公司利用大量音乐收藏来推断品味和偏好的趋势。对这些大型数据集的严格分析需要在计算约束和统计性能之间取得平衡,而将这种权衡操作化涉及算法和基于设计的近似的组合。例如,降低图像的分辨率或对高吞吐量数据进行二次采样可以加快计算速度,但会删除潜在的有价值信息。与此同时,对科学过程的详细解释是可信的,因为它们解释了现实世界的复杂性,但估计复杂的模型需要更多的数据和更大的计算机。拟议的工作调查了现代统计和机器学习方法回答应用科学问题的可行性。该项目将创建新的算法和开源软件,将计算近似与正则化相结合,用于分析大型数据集,并为其统计特性提供理论依据。更全面地了解统计正则化、计算近似和科学简约之间的相互作用,将有助于基础科学的进步。PI将采用本项目中开发的方法,以促进气候科学,生物学,音乐分析,天文学,经济学和化学方面的大型数据集的新科学。此外,PI将仔细整合研究目标与教育和推广目标,以吸引小学和高中音乐学生,向他们介绍现代统计和计算机科学,以及将代表性不足的人群纳入研究。大型数据集的计算易处理性和统计效率需要近似或正则化,这两种方法中的任何一种都在理论上平衡了对数据的保真度与科学目标,如简约性、平滑性、稀疏性或可解释性。PI试图阐明正则化和近似作为更好的科学理解工具的双重作用。目前计算机科学的研究集中在改进算法,使计算具有最小的近似。与此同时,统计学家开发了正则化技术,以利用简单的结构-图拓扑,稀疏线性模型,光滑函数-如果代表真理,将改善推理和预测。挑战在于理解这种耦合对科学可解释性的后果。PI试图解决两个重要的问题(1)当计算非常重要时,我们如何选择调优参数?(2)科学结论的准确性和稳定性如何与用于获得这些结论的近似/正则化方法相关联?该项目旨在通过理解计算近似和统计正则化之间的联系来实现基础科学进步,从而促进改进推断。PI将在更合理的统计假设下,通过推导实用算法以及相应的理论依据来填补这一空白。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Volumes of data that seemed unimaginable only a decade ago are now ubiquitous in scientific research. Investment decisions are based on prices, updated every microsecond, for thousands of securities; atmospheric scientists use multiresolution satellite images to understand climate change; and internet companies exploit massive music collections to infer trends in tastes and preferences. Rigorous analysis of these large datasets requires a balance between computational constraints and statistical performance, and operationalizing such tradeoffs involves a combination of algorithmic and design-based approximations. For example, decreasing the resolution of an image or subsampling high-throughput data enables faster computations but removes potentially valuable information. At the same time, elaborate explanations of the scientific process are credible because they account for real-world complexity, but estimating complex models requires both more data and larger computers. The proposed work investigates the viability of modern statistical and machine learning methodologies for answering applied scientific questions. The project will create new algorithms and open-source software for combining computational approximations with regularization for analyzing large datasets as well as providing theoretical justification for their statistical properties. A more complete picture of the interplay between statistical regularization, computational approximation, and scientific parsimony will enable fundamental scientific advancement. The PI will employ the methodologies developed in this project to facilitate novel science with large datasets in climate science, biology, music analysis, astronomy, economics, and chemistry. Furthermore, the PI will carefully integrate the research aims with educational and outreach objectives to engage elementary and high school music students, introducing them to modern statistics and computer science, as well as integrating underrepresented populations in research.Computational tractability and statistical efficiency for large datasets necessitate approximation or regularization, either of which heuristically balances fidelity to the data with scientific goals like parsimony, smoothness, sparsity, or interpretability. The PI seeks to elucidate the dual roles of regularization and approximation as tools for better scientific understanding. Current research in computer science has focused on improving algorithms so as to enable computation with minimal approximation. Meanwhile, statisticians have developed regularization techniques in order to take advantage of simple structures - graph topologies, sparse linear models, smooth functions - that, if representative of the truth, will improve inference and prediction. The challenge is to understand the consequences of this coupling for scientific interpretability. The PI seeks to address two important issues (1) how do we select tuning parameters when computations are at a premium? and (2) how does the accuracy and stability of scientific conclusions relate to the approximation/regularization methods used to obtain those conclusions? This project seeks to enable fundamental scientific progress by understanding the connections between computational approximations and statistical regularization, thereby facilitating improved inferences. The PI will fill this gap by deriving practical algorithms with accompanying theoretical justification under more reasonable statistical assumptions. These tools will be tightly coupled with applications in neuroscience, genetics, atmospheric science, and music.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1609/aaai.v33i01.3301614
发表时间: 2018-05
期刊:
影响因子: --
作者: [Arash Khodadadi;D. McDonald]
通讯作者: Arash Khodadadi;D. McDonald
DOI: 10.1080/00949655.2018.1491575
发表时间: 2016-02
期刊: Journal of Statistical Computation and Simulation
影响因子: 1.2
作者: [D. Homrighausen;D. McDonald]
通讯作者: D. Homrighausen;D. McDonald
Compressed and Penalized Linear Regression
压缩和惩罚线性回归
DOI: 10.1080/10618600.2019.1660179
发表时间: 2020
期刊: Journal of Computational and Graphical Statistics
影响因子: 2.4
作者: [Homrighausen, Darren, McDonald, Daniel J.]
通讯作者: McDonald, Daniel J.
Collaborative research: Statistical and computational efficiency for massive datasets via approximation-regularization
  • 批准号:
    1407439
  • 项目类别:
    Standard Grant
  • 资助金额:
    $8.99万
  • 财政年份:
    2014
  • 负责人:
    Daniel McDonald
  • 依托单位:
海外基金