High-Dimensional Random Forests Learning, Inference, and Beyond
High-Dimensional Random Forests Learning, Inference, and Beyond
批准号:
2310981
负责人:
Yingying Fan
金额:
$25.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-15 至 2026-07-31
中文摘要
随机森林是最常用的预测计算方法之一。这种方法的工作原理是创建一组决策者,就像一个专家团队,然后汇总这些专家的个人预测,形成最终的预测。随机森林的巨大成功已经被应用于许多不同类型的数据时的上级性能所验证。尽管随机森林算法取得了巨大的成功,但由于理论认识的局限性,它在很大程度上仍然被认为是一种黑箱方法,算法本身的复杂性和理论认识的缺乏也使得它产生的结果重现性差,难以解释。该项目将从理论上研究随机森林的属性,以了解算法何时有效,更重要的是,算法何时失败。这样的研究可以为从业者提供更多的信心和更好的指导应用随机森林。该项目将研究如何提高随机森林的可解释性。最后,根据这些研究中获得的理解,该项目将研究如何提高算法的性能,使其对大数据分析更加有用。这些研究活动将为下一代统计学家和数据科学家的专业发展提供许多培训举措。最近,在随机森林算法的分析方面取得了重要进展,例如,在不对回归函数和特征分布做具体假设的情况下,证明了原始版本的随机森林在高维环境下的多项式一致率。然而,仍然有许多根本性的重要问题没有得到回答。该项目的总体目标是深入了解复杂的集成方法,如随机森林,并提供改进的,可解释的和可重复的统计估计和推断结果。该项目将首先研究有关随机森林的一些重要的开放问题,然后转移到统计推断。特别是,最近的研究已经证实,随机森林可以适应稀疏模型。一个自然的问题是如何破坏潜在的真实稀疏结构。此外,一些初步的结果表明,流行的现有方法是有偏见的,当存在特征共线性。该项目将开发有效的特征重要性度量,并进一步研究在存在特征共线性的情况下评估条件特征重要性的p值的计算。该项目还将超越随机森林,并研究条件独立性测试的更大问题。该项目将利用这些理论研究中获得的见解,进一步开发改进的集成学习方法,以提高大数据分析的预测性、可解释性和可再现性。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Random Forests are one of the most popularly used computational methods for making predictions. The approach works by creating a group of decision-makers, like a team of experts, and then aggregates the individual predictions by these experts to form the final prediction. The great success of Random Forests has been verified by the superior performance when applied to many different types of data. Despite the tremendous success, Random Forests are still largely regarded as a Black-box method because of the limited theoretical understanding of it. The complicated nature of the algorithm and lack of theoretical understanding also make the results it produces less reproducible and hard to interpret. The project will theoretically study the properties of Random Forests to understand when the algorithm works, and more importantly, when the algorithm fails. Such studies can provide practitioners with more confidence and better guidance in applying Random Forests. The project will investigate how to improve the interpretability of Random Forests. Finally, with the understanding gained from these studies, the project will study how to improve the performance of the algorithm to make it even more useful for big data analysis. These research activities will offer numerous training initiatives for professional development of the next generation of statisticians and data scientists.Recently, there has been made important progress in the analysis of random forest algorithms, for instance, proof of the polynomial consistency rate of the original version of Random Forests in the high dimensional setting, without making specific assumptions of the regression function and feature distribution. Yet, there are still many fundamentally important questions left unanswered. The overall objective of this project is to provide an in-depth understanding of complicated ensemble methods such as Random Forests, and provide improved, interpretable, and reproducible statistical estimation and inference results. The project will first study some important open questions about Random Forests, and then move to the statistical inference. In particular, recent studies have confirmed that Random Forests can adapt to sparse models. A natural question is how to undermine the underlying true sparsity structure. Furthermore, some preliminary results suggest that popular existing methods are biased when there exists feature collinearity. The project will develop valid feature importance measures and further investigate the calculation of p-values for evaluating conditional feature importance in the existence of feature collinearity. The project will also move beyond Random Forests and study the larger problem of the conditional independence test. Utilizing the insights gained from these theoretical studies, the project will further develop an improved ensemble learning method for better prediction, interpretability, and reproducibility in big data analysis.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FRG: Collaborative Research: Flexible Network Inference
-
批准号:2052964
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2021
-
负责人:Yingying Fan
-
依托单位:
CAREER: High-Dimensional Variable Selection in Nonlinear Models and Classification with Correlated Data
-
批准号:1150318
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2012
-
负责人:Yingying Fan
-
依托单位:
Regularization Methods in High Dimensions with Applications to Functional Data Analysis, Mixed Effects Models and Classification
-
批准号:0906784
-
项目类别:Continuing Grant
-
资助金额:$20.08万
-
财政年份:2009
-
负责人:Yingying Fan
-
依托单位:
海外基金