课题基金 / 基金详情

Shrinkage Methods for Variable Selection and Structure Discovery, with Applications to High Dimensional Data

Shrinkage Methods for Variable Selection and Structure Discovery, with Applications to High Dimensional Data
用于变量选择和结构发现的收缩方法及其在高维数据中的应用
批准号:
1005612
负责人:
Howard Bondell
金额:
$13.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-08-15 至 2014-07-31

项目摘要

项目成果

Howard Bondell的其他基金

相似基金

相关文献

中文摘要
翻译
拟议的研究解决了当今复杂数据结构背景下的变量选择问题。该项目的第一个组成部分是将“监督聚类”和变量选择结合到一个步骤中。其目标是促进识别重要的预测性集群,从而发现潜在的分组结构。其次,本项目提出了一种调整变量选择方法的新技术,该技术不仅性能优于现有方法,而且对非统计学家具有直观的解释。这个项目的第三个组成部分是执行同时变量选择和约束分位数回归。在复杂数据中,简单地将平均值建模为预测值的函数可能不能捕捉到全部关系。分位数回归模拟预测因素对反应的不同百分位数的影响。本项目中提出的方法缓解了分位数回归中众所周知的曲线交叉问题。最后,在复杂和高维数据中,离群点几乎肯定会存在,因此对这些离群点具有健壮性的方法至关重要。该项目的最后一个组成部分是一种将稳健方法与通过直接加权观测进行变量选择相结合的方法。这项研究的所有四个组成部分都将通过惩罚,或者等价地,通过适当选择惩罚的约束优化框架来发展。随着现在所有科学领域都有丰富的信息可用,决定将大量可能的预测变量中的哪些包括在模型中可能是一项压倒性的任务。因此,开发执行变量选择的技术是至关重要的。进行变量选择的惩罚技术越来越受欢迎,并经常应用于主题研究的不同分支,包括药物发现、消费者营销、环境系统、金融市场、图像处理、国土安全、基因组学、蛋白质组学和代谢组学。通常情况下,调查人员在执行统计分析时会考虑到一些目标。一个常见的例子出现在基因表达研究中,其中人们可能希望同时执行主题分类、基因选择和基因聚类。拟议的研究特别着眼于实现这些类型的多方面分析。拟议研究的一个总主题是,适当选择惩罚函数可以同时以综合方式实现多个统计目标。变量选择问题在所有学科中的重要性,以及研究人员S与医学研究人员和其他科学家的合作,将使结果容易地传播到应用研究社区,在那里它可以用于改善生活质量。
英文摘要
The proposed research addresses variable selection within the context of today's complex data structures. The first component of this project is to combine "supervised clustering" and variable selection into a single step. The goal is to facilitate the identification of important predictive clusters leading to the discovery of an underlying grouping structure. Secondly, this project proposes to introduce a new technique to tune variable selection methods which not only outperforms existing methods, but has an intuitive interpretation to non-statisticians. The third component of this project is to perform simultaneous variable selection and constrained quantile regression. In complex data, simply modeling the mean as a function of the predictors may not capture the full relationships. Quantile regression models the effect of the predictors on various percentiles of the response. The approach proposed in this project alleviates the well-known issue of crossing curves in quantile regression. Finally, in complex and high-dimensional data, outliers are almost certain to exist, so methods that are robust to these outliers are essential. The final component of this project is an approach to combine robust methods with variable selection via a direct weighting of observations. All four components of this research will be developed via a penalization, or equivalently, a constrained optimization, framework via appropriate choices of penalty. With the abundance of information now available in all scientific fields, it can be an overwhelming task to decide on which of the massive number of possible predictor variables to include in a model. Therefore, it is essential to develop techniques to perform variable selection. Penalization techniques to perform variable selection have gained increasing popularity and are routinely applied in diverse branches of subject-matter research including drug discovery, consumer marketing, environmental systems, financial markets, image processing, homeland security, genomics, proteomics, and metabolomics. It is often the case that the investigator has a number of goals in mind when performing a statistical analysis. A common example occurs in gene expression studies, where one may wish to perform subject classification, gene selection, and gene clustering, simultaneously. The proposed research is particularly geared toward enabling the accomplishment of these types of multi-faceted analyses. A general theme of the proposed research is that appropriately chosen penalty functions can achieve multiple statistical objectives simultaneously and in an integrated fashion. The importance of the variable selection problem across all disciplines, and the investigator?s collaborations with medical researchers and other scientists will allow the results to be readily disseminated into the applied research community where it can be used to improve the quality of life.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Comprehensive Framework for Fully Efficient Robust Estimation and Variable Selection, with Application to High-Dimensional and Complex Data
  • 批准号:
    1308400
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2013
  • 负责人:
    Howard Bondell
  • 依托单位:
Advances in Variable Selection with Grouped Predictors
  • 批准号:
    0705968
  • 项目类别:
    Standard Grant
  • 资助金额:
    $14.0万
  • 财政年份:
    2007
  • 负责人:
    Howard Bondell
  • 依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data