High-Dimensional Random Forests Learning, Inference, and Beyond
High-Dimensional Random Forests Learning, Inference, and Beyond
批准号:
2310981
负责人:
Yingying Fan
金额:
$25.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-15 至 2026-07-31
中文摘要
随机森林是最常用的预测计算方法之一。这种方法的工作方式是创建一组决策者,比如一个专家团队,然后将这些专家的个人预测汇总起来,形成最终的预测。随机森林的巨大成功已经被应用于许多不同类型的数据时的优越性能所验证。尽管随机森林方法取得了巨大的成功,但由于理论上对它的理解有限,它在很大程度上仍被认为是一种黑箱方法。该算法的复杂性和缺乏理论理解也使其产生的结果较难重复性和难以解释。该项目将从理论上研究随机森林的属性,以了解算法何时起作用,更重要的是,算法何时失败。这样的研究可以为从业者在应用随机森林时提供更多的信心和更好的指导。该项目将研究如何提高随机森林的可解释性。最后,随着从这些研究中获得的理解,该项目将研究如何提高算法的性能,使其更适用于大数据分析。这些研究活动将为下一代统计学家和数据科学家的专业发展提供许多培训举措。最近,在随机森林算法的分析方面取得了重要进展,例如证明了原始版本的随机森林在高维环境下的多项式一致性,而不对回归函数和特征分布做出具体假设。然而,仍有许多根本性的重要问题悬而未决。这个项目的总体目标是深入了解复杂的集合方法,如随机森林,并提供改进的、可解释的和可重复的统计估计和推断结果。该项目将首先研究关于随机森林的一些重要的公开问题,然后转移到统计推断。特别是,最近的研究证实了随机森林可以适应稀疏模型。一个自然的问题是,如何破坏潜在的真正稀疏性结构。此外,一些初步的结果表明,当存在特征共线时,现有的流行方法是有偏差的。该项目将开发有效的特征重要性度量,并进一步研究在存在特征共线的情况下用于评估条件特征重要性的p值的计算。该项目还将超越随机森林,研究条件独立性测试这一更大的问题。利用从这些理论研究中获得的见解,该项目将进一步开发改进的集成学习方法,以在大数据分析中更好地预测、解释和重复性。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Random Forests are one of the most popularly used computational methods for making predictions. The approach works by creating a group of decision-makers, like a team of experts, and then aggregates the individual predictions by these experts to form the final prediction. The great success of Random Forests has been verified by the superior performance when applied to many different types of data. Despite the tremendous success, Random Forests are still largely regarded as a Black-box method because of the limited theoretical understanding of it. The complicated nature of the algorithm and lack of theoretical understanding also make the results it produces less reproducible and hard to interpret. The project will theoretically study the properties of Random Forests to understand when the algorithm works, and more importantly, when the algorithm fails. Such studies can provide practitioners with more confidence and better guidance in applying Random Forests. The project will investigate how to improve the interpretability of Random Forests. Finally, with the understanding gained from these studies, the project will study how to improve the performance of the algorithm to make it even more useful for big data analysis. These research activities will offer numerous training initiatives for professional development of the next generation of statisticians and data scientists.Recently, there has been made important progress in the analysis of random forest algorithms, for instance, proof of the polynomial consistency rate of the original version of Random Forests in the high dimensional setting, without making specific assumptions of the regression function and feature distribution. Yet, there are still many fundamentally important questions left unanswered. The overall objective of this project is to provide an in-depth understanding of complicated ensemble methods such as Random Forests, and provide improved, interpretable, and reproducible statistical estimation and inference results. The project will first study some important open questions about Random Forests, and then move to the statistical inference. In particular, recent studies have confirmed that Random Forests can adapt to sparse models. A natural question is how to undermine the underlying true sparsity structure. Furthermore, some preliminary results suggest that popular existing methods are biased when there exists feature collinearity. The project will develop valid feature importance measures and further investigate the calculation of p-values for evaluating conditional feature importance in the existence of feature collinearity. The project will also move beyond Random Forests and study the larger problem of the conditional independence test. Utilizing the insights gained from these theoretical studies, the project will further develop an improved ensemble learning method for better prediction, interpretability, and reproducibility in big data analysis.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FRG: Collaborative Research: Flexible Network Inference
-
批准号:2052964
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2021
-
负责人:Yingying Fan
-
依托单位:
CAREER: High-Dimensional Variable Selection in Nonlinear Models and Classification with Correlated Data
-
批准号:1150318
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2012
-
负责人:Yingying Fan
-
依托单位:
Regularization Methods in High Dimensions with Applications to Functional Data Analysis, Mixed Effects Models and Classification
-
批准号:0906784
-
项目类别:Continuing Grant
-
资助金额:$20.08万
-
财政年份:2009
-
负责人:Yingying Fan
-
依托单位:
海外基金