Leveraging Structural Information in Regression Tree Ensembles
Leveraging Structural Information in Regression Tree Ensembles
批准号:
2015636
负责人:
Antonio Linero
金额:
$2.59万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-01 至 2020-08-31
中文摘要
统计学中的一项常见任务是预测;例如,从业者可能有兴趣在给定关于个人的遗传信息的情况下预测疾病的存在。由于最近在数据收集方面的进步,人们经常可以访问包含大量预测者的数据集,但相应地只有很少的对象。这种设置通常被称为“大P,小n”方案。在这种情况下得出有意义的结论通常是不可能的,除非基础数据满足某些结构性假设。最简单的这种结构性假设是,只有一小部分预测因素是相关的;在这种情况下,找到有用的预测因素就像是在大海捞针。这个项目的目标是构建适应这一结构假设和其他结构假设的程序。该项目将专注于基于决策树的方法,这是一种类似流程图的结构,其中预测基于预测者是否满足各种规则。通常会构建决策树的集成,并对每棵树的预测进行平均。虽然决策树集成经常用于高维数据,但尚不清楚它们在多大程度上适应数据的结构属性。该项目将表明,在实践中,现有的决策树集成方法不适合常见的结构假设,并将开发新的方法。除了开发具有强大理论支持的方法外,该项目还将支持R包的开发,使从业者能够轻松地使用我们的方法。PI将开发贝叶斯方法,将结构信息整合到基于树的集成方法中,并从理论上确立利用这些额外信息的好处。这形成了线性模型中使用的参数方法的非参数对应,例如套索、图形套索或组合套索;参数设置中的贝叶斯方法包括使用变量选择先验,例如钉子和板材先验和全局局部收缩先验。结构信息将通过修改决策树集成上常用的先验来结合,从而使先验集中在满足期望结构的模型上。PI将首先研究稀疏性诱导先验的理论性质,该先验旨在消除不必要的预测因素。这里的稀疏性是通过在给定分支与给定预测器相关联的先验概率之前应用稀疏性诱导Dirichlet来获得的。该先验将被扩展以允许以类似于组套索的方式通过考虑Dirichlet树先验来进行分组变量选择,并且进一步通过稀疏性诱导逻辑正态先验来适应预测器中的图形结构。此外,PI将开发计算效率高的马尔可夫链蒙特卡罗算法,以适应所产生的模型。与现有方法相比,这些结构先验将被证明导致预测精度的大幅提高,并导致更准确的科学发现。
英文摘要
A common task in statistics is prediction; for example, a practitioner may be interested in predicting the presence of a disease given genetic information about an individual. Due to recent advances in data collection, frequently one has access to datasets which contain a massive number of predictors, but with correspondingly few subjects. This setting is generally referred to as the "big P, small n" scenario. Drawing meaningful conclusions under such circumstances is generally impossible unless the underlying data satisfy certain structural assumptions. The simplest such structural assumption is that only a small number of the predictors are relevant; in this setting, finding the useful predictors corresponds to finding a so-called "needle in a haystack." The goal of this project is to construct procedures which adapt to this, and other, structural assumptions. The project will focus on methods based on decision trees, which are flowchart-like structures in which predictions are based on whether the predictors satisfy various rules. Usually an ensemble of decision trees are constructed, with the predictions for each individual tree averaged. While decision tree ensembles are frequently used with high dimensional data, it is unclear to what extent they adapt to the structural properties of the data. This project will show that, in practice, off-the-shelf decision tree ensembling methods do not adapt to common structural assumptions, and will develop new methods which do. In addition to developing methods with strong theoretical support, this project will support the development of an R package to give practitioners easy access to our methodology. The PI will develop Bayesian methods for incorporating structural information into tree-based ensemble methods, and establish theoretically the benefit of making use of this additional information. This forms a nonparametric counterpart to the parametric approaches used in linear models, such as the lasso, graphical lasso, or group lasso; Bayesian approaches in the parametric setting include the use of variable selection priors, such as spike-and-slab priors and global-local shrinkage priors. Structural information will be incorporated by modifying the commonly used priors on decision tree ensembles so that the prior is concentrated on models which satisfy the desired structure. The PI will first investigate the theoretical properties of a sparsity inducing prior which is designed to eliminate unnecessary predictors. Sparsity here is obtained by applying a sparsity inducing Dirichlet prior to the a priori probability that a given branch is associated to a given predictor. This prior will be extended to allow for grouped variable selection in a similar manner to the group lassoby considering the class of Dirichlet tree priors, and further to accommodate graphical structures in the predictors through sparsity inducing logistic normal priors. Additionally, the PI will develop computationally efficient Markov chain Monte Carlo algorithms to fit the resulting models. Compared to existing methods, these structural priors will be shown to lead to substantial gains in predictive accuracy, and to more accurate scientific discovery.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1080/00224065.2020.1801366
发表时间:
2020
期刊:
Journal of Quality Technology
影响因子:
2.5
作者:
[Shamp, Wright, Varbanov, Roumen, Chicken, Eric, Linero, Antonio, Yang, Yun]
通讯作者:
Yang, Yun
Semiparametric mixed‐scale models using shared Bayesian forests
使用共享贝叶斯森林的半参数混合尺度模型
DOI:
10.1111/biom.13107
发表时间:
2019
期刊:
Biometrics
影响因子:
1.9
作者:
[Linero, Antonio R., Sinha, Debajyoti, Lipsitz, Stuart R.]
通讯作者:
Lipsitz, Stuart R.
CAREER: Foundations for Bayesian Nonparametric Causal Inference
-
批准号:2144933
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2022
-
负责人:Antonio Linero
-
依托单位:
Leveraging Structural Information in Regression Tree Ensembles
-
批准号:1712870
-
项目类别:Continuing Grant
-
资助金额:$10.0万
-
财政年份:2017
-
负责人:Antonio Linero
-
依托单位:
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位: