课题基金 / 基金详情

New algorithms for consistent model selection beyond linear models

New algorithms for consistent model selection beyond linear models
用于超越线性模型的一致模型选择的新算法
批准号:
1607840
负责人:
Xuming He
金额:
$30.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2020-08-31

项目摘要

项目成果

Xuming He的其他基金

相似基金

相关文献

中文摘要
翻译
统计模型的建立是科学发现的重要组成部分。在大数据时代,高维数据频繁出现。在线性模型、广义线性模型和删减数据模型的框架下,高维特征的模型选择是近年来研究的一个非常活跃的领域。PI的目标是在贝叶斯计算框架内开发新的模型选择算法,该算法可扩展到高维问题。PI通过与大气科学、遗传学和运动机能学方面的科学家合作来推动拟议的研究,旨在开发广泛适用于统计建模和数据分析的方法。最近的大部分工作都集中在通过惩罚或规范化来缩小规模。广义的贝叶斯计算方法在统计学中发挥着重要的作用,包括模型选择和估计,但在高维统计中面临着重要的障碍,无论是在理论复杂性还是在计算可扩展性方面。PI旨在开发一个理论框架,从频率论的角度来证明模型选择的一致性,这为为什么贝叶斯模型选择方法可以提供L0惩罚的渐近逼近提供了有趣的见解。提出的工作的一个重要部分是在稀疏模型的选择中改进Gibbs采样器的开发,该采样器在存在高维变量时比标准MCMC算法更具可扩展性。贝叶斯方法在非凸目标函数问题中特别有用,贝叶斯计算方法在性能上比直接优化更健壮。该项目中考虑的此类问题的主要应用是对删减数据进行分位数回归。除了模型选择之外,PI还提出了一种新的估计方法,用于审查分位数回归,该方法有望在计算和统计上高效。同样重要的是,新方法很容易适应其他估计方法难以处理的一般形式的审查。PI将继续与博士生合作,并为本科生提供研究经验,将研究与教育结合起来。研究成果将通过会议和讲习班以及通过在广泛阅读的统计科学期刊上发表来适当传播。
英文摘要
Statistical model building is an important part of scientific discovery. In the big data era, high dimensional data arise frequently. Model selection in the presence of high dimensional features in the framework of linear models, generalized linear models, and models with censored data has been a very active area of research in recent years. The PI aims to develop new algorithms for model selection, within a Bayesian computational framework, that are scalable for high dimensional problems. The PI motivates the proposed research through collaborations with scientists in atmospheric sciences, genetics, and kinesiology, and aims to develop methodologies that are broadly applicable in statistical modeling and data analysis. Much of the recent work has focused on shrinkage through penalization or regularization. Bayesian computational methods, when interpreted broadly, play a valuable role in statistics, including model selection and estimation, but face important hurdles in high dimensional statistics, both in theoretical intricacy and in computational scalability. The PI aims to develop a theoretical framework to demonstrate model selection consistency from the frequentist perspective, which offers interesting insights into why Bayesian model selection methods can provide an asymptotic approximation to the L0 penalty. An important part of the proposed work is the development of a modified Gibbs sampler in the selection of sparse models that is much more scalable than standard MCMC algorithms in the presence of high dimensional variables. The Bayesian methods are especially useful in problems with non-convex objective functions, where Bayesian computation methods can be more robust in performance than direct optimization. A primary application of such a problem considered in the project is quantile regression for censored data. In addition to model selection, the PI proposes a new estimation method for censored quantile regression that promises to be computationally and statistically efficient. Equally importantly, the new method adapts easily to general forms of censoring that other estimation methods have found difficult to handle. The PI will continue integrating research with education by working with PhD students and by providing research experiences for undergraduate students. The research output will be properly disseminated through conferences and workshops and through publication in widely read journals in statistical science.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Conference: Workshop on Translational Research on Data Heterogeneity
  • 批准号:
    2406154
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.6万
  • 财政年份:
    2024
  • 负责人:
    Xuming He
  • 依托单位:
Covariate-adjusted Expected Shortfall under Data Heterogeneity
  • 批准号:
    2345035
  • 项目类别:
    Standard Grant
  • 资助金额:
    $33.0万
  • 财政年份:
    2023
  • 负责人:
    Xuming He
  • 依托单位:
Covariate-adjusted Expected Shortfall under Data Heterogeneity
Towards Efficient Bias Correction in Data Snooping
国内基金
海外基金
固定参数可解算法在平面图问题的应用以及和整数线性规划的关系
  • 批准号:
    60973026
  • 项目类别:
    面上项目
  • 资助金额:
    32.0万元
  • 批准年份:
    2009
  • 负责人:
    鲁道夫
  • 依托单位:
Computational Methods for Analyzing Toponome Data