课题基金 / 基金详情

Advancing Theory and Methodology for Tree-Based Algorithms in High Dimensions

Advancing Theory and Methodology for Tree-Based Algorithms in High Dimensions
推进高维树基算法的理论和方法
批准号:
2209975
负责人:
Bin Yu
金额:
$33.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-15 至 2025-06-30

项目摘要

项目成果

Bin Yu的其他基金

相似基金

相关文献

中文摘要
翻译
预测统计建模长期以来一直是科学和工程的支柱。近年来,大数据的激增导致需要超越传统的线性模型,并需要能够利用复杂非线性关系的灵活模型。基于决策树的模型已经成为一种易于使用和高性能的模型,特别是对于非结构化的表格数据集,如电子健康记录,它们通常优于神经网络。此外,由于决策树可以很容易地被非专家可视化和模拟,这使得它们比黑箱机器学习模型更容易审计,这在预测用于指导诊所或法庭上的高风险决策时尤为重要。不幸的是,基于决策树的模型在统计学上并没有得到很好的理解,而且仍然不清楚各种模型何时以及为什么获得更好的相对预测性能。该项目计划通过识别真实的世界数据集中的结构属性来弥合这一差距,这些结构属性使它们适合或不适合当前基于树的模型。然后,这种理解将用于开发基于决策树的更好的算法,以及从这些模型中提取可重复的科学见解的方法。在项目期间,研究生将接受理论、领域驱动的数据科学和开源软件开发方面的培训。研究成果将通过课程、即将出版的书籍以及在研讨会和会议上的演讲进一步传播。该项目计划两个重点来发展决策树和随机森林的相关理论。首先,它将分析基于树的算法在一系列不同的生成回归模型上的泛化性能,以得出它们的归纳偏差。归纳偏差是机器学习中的一个众所周知的概念,它被定义为算法在推广到新数据时所做的假设。由于真实的数据集通常呈现出一些可以使用正确的归纳偏差来利用的结构,因此该项目的结果将允许更好地识别在给定应用中选择哪种算法,从而改进决策树和随机森林的经典非参数回归分析。其次,该项目将研究一个新的一般框架,用于使用杂质平均减少(MDI)特征重要性来获得模型无关的非线性特征重要性度量。这个框架利用了MDI的一个新的解释,从线性回归的r平方值,是渐近有效的,即使用于生成MDI的决策树不一定是一个很好的模型为基础的回归function.This奖项反映了NSF的法定使命,并已被认为是值得通过使用基金会的智力价值和更广泛的影响审查标准进行评估的支持。
英文摘要
Predictive statistical modeling has long been part of the backbone of science and engineering. In recent years, the proliferation of big data has led to a need to go beyond traditional linear models, and a need for flexible models that can exploit complicated nonlinear relationships. Models based on decision trees have emerged as an easy-to-use and high performing class of models, especially for unstructured tabular datasets such as electronic health records, in which they have been found to typically outperform neural networks. Furthermore, since decision trees can be easily visualized and simulated by non-experts, this makes them easier to audit than black box machine learning models, which is especially important when predictions are used to guide high-stakes decisions in the clinic or the courtroom. Unfortunately, models based on decision trees are not well understood statistically, and it is still unclear when and why various models obtain better relative predictive performance. The project plans to bridge this gap by identifying structural properties in real world datasets that make them either amenable or not amenable to current tree-based models. This understanding will then be used to develop better algorithms based on decision trees, as well as methodology to extract reproducible scientific insights from these models. In the duration of the project, graduate students will be trained in theory, domain-driven data science, and open-source software development. Research results will further be disseminated through courses, an upcoming book, and presentations at workshops and conferences.The project plans two thrusts to develop relevant theory for decision trees and random forests. First, it will analyze the generalization performance of tree-based algorithms on a range of different generative regression models in order to elicit their inductive bias. Inductive bias is a well-known concept from machine learning, and is defined as the assumptions an algorithm makes when generalizing to new data. Since real world datasets often present some structure that can be exploited using the right inductive bias, results of this project will allow better identification of which algorithm to choose in a given application, thus improving on classical nonparametric regression analysis of decision trees and random forests. Second, the project will study a new general framework for obtaining model-agnostic nonlinear feature significance measures using mean decrease in impurity (MDI) feature importance. This framework makes use of a novel interpretation of MDI in terms of r-squared values from linear regression, and is asymptotically valid even if the decision tree used to generate MDI is not necessarily a good model for the underlying regression function.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Understanding Complexity and the Bias-Variance Tradeoff in High Dimensions: Theory and Data Evidence
  • 批准号:
    2015341
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2020
  • 负责人:
    Bin Yu
  • 依托单位:
Parallel Ensemble Learning and Feature Interaction Discovery: High Volume Dynamic Data
  • 批准号:
    1953191
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.2万
  • 财政年份:
    2020
  • 负责人:
    Bin Yu
  • 依托单位:
Understand the functional mechanism of the DSP1 complex in the 3' end maturation of plant small nuclear RNAs
  • 批准号:
    1818082
  • 项目类别:
    Standard Grant
  • 资助金额:
    $68.26万
  • 财政年份:
    2018
  • 负责人:
    Bin Yu
  • 依托单位:
BIGDATA: F: Scalable and Interpretable Machine Learning: Bridging Mechanistic and Data-Driven Modeling in the Biological Sciences
  • 批准号:
    1741340
  • 项目类别:
    Standard Grant
  • 资助金额:
    $90.0万
  • 财政年份:
    2017
  • 负责人:
    Bin Yu
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
基于isomorph theory研究尘埃等离子体物理量的微观动力学机制
  • 批准号:
    12247163
  • 项目类别:
    专项项目
  • 资助金额:
    18.00万元
  • 批准年份:
    2022
  • 负责人:
    黄栋
  • 依托单位:
Toward a general theory of intermittent aeolian and fluvial nonsuspended sediment transport
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    55万元
  • 批准年份:
    2022
  • 负责人:
    Thomas Pahtz
  • 依托单位:
英文专著《FRACTIONAL INTEGRALS AND DERIVATIVES: Theory and Applications》的翻译
  • 批准号:
    12126512
  • 项目类别:
    数学天元基金项目
  • 资助金额:
    12.0万元
  • 批准年份:
    2021
  • 负责人:
    李常品
  • 依托单位: