课题基金 / 基金详情

Collaborative research: Statistical and computational efficiency for massive datasets via approximation-regularization

Collaborative research: Statistical and computational efficiency for massive datasets via approximation-regularization
协作研究:通过近似正则化实现海量数据集的统计和计算效率
批准号:
1407439
负责人:
Daniel McDonald
金额:
$8.99万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2018-08-31

项目摘要

项目成果

Daniel McDonald的其他基金

相似基金

相关文献

中文摘要
翻译
该项目将计算机科学的近似方法与现代统计理论相结合,以改进对大数据集的分析。现代统计分析需要在大型数据集上计算可行的方法,同时保持统计效率。通常,这两个问题被认为是矛盾的:相对于精确方法,允许计算的近似方法被认为会降低统计性能。从统计学的角度来看,精确解是不可取的,而正则化解是首选。正则化可以被认为是在数据保真度和遵守有关数据生成过程的先验知识(例如平滑性或稀疏性)之间进行形式化的权衡。得到的估计器往往更有用、更可解释,并且适合作为其他方法的输入。相反,在计算机科学应用中,当前大部分关于近似方法的工作都在其中,输入通常被认为是被精确观察到的。普遍的理念是,虽然遗憾的是,确切的问题无法解决,但任何近似解都应该尽可能接近精确解。我们做出了一个重要的认识:近似方法本身自然会导致正则化,这表明一些计算近似可以同时实现海量数据分析,同时增强统计性能的有趣可能性。我们的研究开发了利用这种现象的新方法,我们将其称为“近似正则化”。第一种方法使用矩阵预处理器来稳定最小二乘准则。如果正确校准,这种方法比正则化最小二乘法具有计算和存储优势,同时提供统计上优越的解决方案。第二项创新解决了大型数据集回归的主成分分析 (PCA) 问题,其中 PCA 在计算上不可行,而且已知在统计上不一致。通过采用随机近似,我们可以解决这两个问题,同时改进预测。最后,我们引入了无监督降维的新方法,利用稀疏性的近似算法和诱导稀疏性的统计方法,使得能够在非常大的矩阵上使用谱技术。在每种情况下,相对于现有方法,近似正则化都会产生计算和统计增益。这项研究认识到近似是正则化,因此可以在实现计算的同时提高统计准确性。它将产生针对大型数据集的新统计方法,这些方法在计算和统计上都优于现有方法,同时也引起人们对统计学中这一重要领域的关注。此外,这些方法将使天文学、遗传学、文本和图像处理、气候科学和预测等其他领域的科学家能够随时利用现有数据。
英文摘要
This project integrates approximation methodology from computer science with modern statistical theory to improve analysis of large data sets. Modern statistical analysis requires methods that are computationally feasible on large datasets while at the same time preserving statistical efficiency. Frequently, these two concerns are seen as contradictory: approximation methods that enable computation are assumed to degrade statistical performance relative to exact methods. The statistical perspective is that the exact solution is undesirable, and a regularized solution is preferred. Regularization can be thought of as formalizing a trade-off between fidelity to the data and adherence to prior knowledge about the data-generating process such as smoothness or sparsity. The resulting estimator tends to be more useful, interpretable, and suitable as an input to other methods. Conversely, in computer science applications, where much of the current work on approximation methods resides, the inputs are generally considered to be observed exactly. The prevailing philosophy is that while the exact problem is, regrettably, unsolvable, any approximate solution should be as close as possible to the exact one. We make a crucial realization: that the approximation methods themselves naturally lead to regularization, suggesting the intriguing possibility that some computational approximations can simultaneously enable the analysis of massive data while enhancing statistical performance. Our research develops new methods that leverage this phenomenon, which we have dubbed 'approximation-regularization.' The first method uses a matrix pre-conditioner to stabilize the least-squares criterion. If properly calibrated, this approach provides computational and storage advantages over regularized least squares while providing a statistically superior solution. A second innovation addresses principal components analysis (PCA) for regression on large data sets where PCA is both computationally infeasible and known to be statistically inconsistent. By employing randomized approximations, we can address both of these issues, while improving predictions at the same time. Lastly, we introduce new methods for unsupervised dimension reduction, whereby approximation algorithms that leverage sparsity, and statistical methods that induce it, enable the use of spectral techniques on very large matrices. In each of these cases, approximation-regularization yields both computational and statistical gains relative to existing methodologies. This research recognizes that approximation is regularization and can thereby increase statistical accuracy while enabling computation. It will result in new statistical methods for large datasets, which are computationally and statistically preferable to existing approaches, while also bringing attention to this important area in statistics. Additionally, these methods will permit scientists in other fields, such as astronomy, genetics, text and image processing, climate science, and forecasting, to make ready use of available data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Calibrating Regularization for Enhanced Statistical Inference
  • 批准号:
    1753171
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2018
  • 负责人:
    Daniel McDonald
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
HIF-1α调控软骨细胞衰老在骨关节炎进展中的作用及机制研究
  • 批准号:
    82371603
  • 项目类别:
    面上项目
  • 资助金额:
    49.00万元
  • 批准年份:
    2023
  • 负责人:
    陈晓
  • 依托单位:
超声驱动压电效应激活门控离子通道促眼眶膜内成骨的作用及机制研究
  • 批准号:
    82371103
  • 项目类别:
    面上项目
  • 资助金额:
    49.00万元
  • 批准年份:
    2023
  • 负责人:
    阮静
  • 依托单位:
Lienard系统的不变代数曲线、可积性与极限环问题研究
  • 批准号:
    12301200
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30.00万元
  • 批准年份:
    2023
  • 负责人:
    钱欣洁
  • 依托单位: