Understanding Complexity and the Bias-Variance Tradeoff in High Dimensions: Theory and Data Evidence
Understanding Complexity and the Bias-Variance Tradeoff in High Dimensions: Theory and Data Evidence
批准号:
2015341
负责人:
Bin Yu
金额:
$30.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-07-01 至 2024-06-30
中文摘要
在过去的十年中,大型机器学习模型在现代数据问题中的使用显著增加;这些模型已经在各种任务中取得了成功,例如图像分类、语言翻译和语音识别。最近,机器学习正在进入新的领域,如机器人、自动驾驶和医学。然而,这些模型通常对扰动不具有鲁棒性,并且容易受到对手的攻击。这些缺点保证了对这些模型的“黑箱”本质的迫切而深刻的理解。主要研究者计划通过以技术方式描述它们的“复杂性”来理解这些模型。基于最小描述长度原则的一种新的复杂性度量,揭示了对经典统计基础的洞察,并告知这些新的高维模型将如何以及何时工作。这种新颖的复杂性度量有望通过样本大小计算和改进有限数据的模型选择,将应用于精确医疗等任务关键领域,在这些领域中,收集标记数据是昂贵的。该研究在统计学和机器学习(包括深度学习)领域具有理论和应用影响。在项目期间,研究生将接受理论、领域驱动数据科学和开源软件开发方面的培训。这项研究将通过课程、一本即将出版的书以及在讲习班和会议上的演讲进一步传播。深度神经网络(DNN)在许多情况下泛化得很好,因为在一个任务上训练的DNN通常在相同任务的类似未见数据上表现良好。它们可以在高度参数化的情况下做到这一点,即参数的数量远远大于训练样本的数量。奥卡姆剃刀和偏差-方差权衡智慧建议,当从具有相似性能的不同复杂性的模型中选择时,更倾向于选择一个简单的模型。深度神经网络的良好性能,尽管过度参数化,导致许多研究人员质疑偏差-方差权衡的经典统计原理(更喜欢一个简单的模型)在现代机器学习(ML)和统计任务中常见的高维设置的有效性。在这个项目中,首席研究员首先重新考虑高维模型的有效复杂性度量的定义——这是奥卡姆剃刀和偏差-方差权衡原则的基础。为高维模型找到一个这样的度量仍然是一项艰巨的任务。仅仅计算参数的数量并不是一个有效的复杂性度量,特别是当训练样本的数量很少的时候。最小描述长度原则将用于提供一种系统的方法来理解高维线性模型、核方法和dnn的复杂性。复杂性度量将作为理解关键概念的基础,例如偏差-方差权衡,以及对高维模型的进一步分析。理论结果将通过一套广泛的数据启发实验来增强。在与新的复杂性度量建立偏差-方差权衡之后,将对这些度量进行调查:(i)从一组竞争模型中选择一个简单模型,其中简单将通过基于mdl的复杂性而不是参数数量来定义,以及(ii)正则化或修剪大型(预训练)模型,例如,在具有有限数据集的迁移学习设置中,通过权衡训练性能和模型的复杂性。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The past decade has witnessed a significant rise in the usage of very large machine-learning models in modern data problems; these models have shown success in a variety of tasks, such as image classification, language translation, and speech recognition. More recently, machine learning is entering new fields, such as robotics, autonomous driving, and medicine. However, these models are often not robust to perturbations and are vulnerable to attacks by adversaries. These shortcomings warrant an urgent and insightful understanding of the "black-box" nature of these models. The principal investigator plans to understand these models by characterizing their "complexity" in a technical manner. A new complexity measure, based on the principle of minimum description length, sheds insight into classical statistical foundations as well as informing how and when these new high-dimensional models will work. This novel complexity measure is promising to enable applications to mission-critical fields like precision medicine, where the collection of a labeled dataset is expensive, by sample-size calculations and improving model selection with limited data. This research has both theoretical and applied impacts in the fields of statistics and machine learning including deep learning. In the duration of the project, graduate students will be trained in theory, domain-driven data science, and open-source software development. The research will be further disseminated through courses, an upcoming book, and presentations at workshops and conferences.Deep neural networks (DNNs) in many cases generalize well in the sense that a DNN trained on one task often performs well on similar unseen data for the same task. They can do so despite being highly overparameterized, i.e., the number of parameters is much larger than the number of training samples. Occam's razor and the bias-variance trade-off wisdom suggest to prefer a simple model when choosing from amongst models of varying complexity with similar performance. The good performance of DNNs, despite the overparametrization, has led many researchers to question the validity of the classical statistical principle of bias-variance trade-off (and preferring a simple model) for high-dimensional settings common in modern machine learning (ML) and statistical tasks. In this project, the principal investigator begins by reconsidering the definition of a valid complexity measure – which forms the basis of Occam’s razor and the bias-variance trade-off principle – for high-dimensional models. Finding one such measure for high-dimensional models has remained a difficult task. Merely counting the number of parameters is not a valid complexity measure, especially when the number of training examples is small. The principle of minimum description length will be used to provide a systematic approach to understanding the complexity of high-dimensional linear models, kernel methods, and finally DNNs. The complexity measure will serve as a basis for understanding key concepts such as the bias-variance trade-off and for further analysis into high-dimensional models. The theoretical results will be augmented with an extensive set of data-inspired experiments. After establishing the bias-variance trade-off with the new complexity measures, these measures will then be investigated for (i) selecting a simple model from amongst a set of competitive models, where simple will be defined via the MDL-based complexity and not the number of parameters, and (ii) regularizing or pruning a large (pre-trained) model, for example, in a transfer learning setting with limited dataset, by trading off the training performance with the complexity of the model.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
MDI+: A Flexible Random Forest-Based Feature Importance Framework
MDI:一种灵活的基于随机森林的特征重要性框架
DOI:
--
发表时间:
2023
期刊:
arXivorg
影响因子:
--
作者:
[Agarwal, Abhineet, Kenney, Ana M., Tan, Yan Shuo, Tang, Tiffany M., Yu, Bin]
通讯作者:
Yu, Bin
Fast Interpretable Greedy-Tree Sums (FIGS)
快速可解释的贪婪树和(FIGS)
DOI:
--
发表时间:
2023
期刊:
ArXivorg
影响因子:
--
作者:
[Tan, Yan Shuo, Singh, Chandan, Nasseri, Keyan, Agarwal, Abhineet, Duncan, James, Ronen, Omer, Epland, Matthew, Kornblith, Aaron, Yu, Bin]
通讯作者:
Yu, Bin
DOI:
--
发表时间:
2021
期刊:
ArXivorg
影响因子:
--
作者:
[Nikhil Ghosh, Song Mei]
通讯作者:
Nikhil Ghosh, Song Mei
An investigation into the effects of pre-training data distributions for pathology report classification
预训练数据分布对病理报告分类影响的调查
DOI:
--
发表时间:
2023
期刊:
arXivorg
影响因子:
--
作者:
[Hsu, Aliyah R., Cherapanamjeri, Yeshwanth, Park, Briton, Naumann, Tristan, Odisho Anobel Y., Yu, Bin]
通讯作者:
Yu, Bin
Advancing Theory and Methodology for Tree-Based Algorithms in High Dimensions
-
批准号:2209975
-
项目类别:Standard Grant
-
资助金额:$33.0万
-
财政年份:2022
-
负责人:Bin Yu
-
依托单位:
Parallel Ensemble Learning and Feature Interaction Discovery: High Volume Dynamic Data
-
批准号:1953191
-
项目类别:Standard Grant
-
资助金额:$45.2万
-
财政年份:2020
-
负责人:Bin Yu
-
依托单位:
Understand the functional mechanism of the DSP1 complex in the 3' end maturation of plant small nuclear RNAs
-
批准号:1818082
-
项目类别:Standard Grant
-
资助金额:$68.26万
-
财政年份:2018
-
负责人:Bin Yu
-
依托单位:
BIGDATA: F: Scalable and Interpretable Machine Learning: Bridging Mechanistic and Data-Driven Modeling in the Biological Sciences
-
批准号:1741340
-
项目类别:Standard Grant
-
资助金额:$90.0万
-
财政年份:2017
-
负责人:Bin Yu
-
依托单位:
Canonical Linear Methods and Hierarchical Non-Linear Methods in High-Dimensional Statistics
-
批准号:1613002
-
项目类别:Continuing Grant
-
资助金额:$60.0万
-
财政年份:2016
-
负责人:Bin Yu
-
依托单位:
Smart Nanofabrication via Rational Assembly of Two-Dimensional Heterosystems
-
批准号:1434689
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2014
-
负责人:Bin Yu
-
依托单位:
Collaborative Research: Leverage Subsampling for Regression and Dimension Reduction
-
批准号:1228246
-
项目类别:Standard Grant
-
资助金额:$22.5万
-
财政年份:2012
-
负责人:Bin Yu
-
依托单位:
Direct Self-Assembly of Large Area, High Crystallinity 2D Graphene on Insulator: An Integratable Carbon Platform
-
批准号:1162312
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2012
-
负责人:Bin Yu
-
依托单位:
Understanding DAWDLE Function in miRNA and siRNA Biogenesis
-
批准号:1121193
-
项目类别:Continuing Grant
-
资助金额:$49.95万
-
财政年份:2011
-
负责人:Bin Yu
-
依托单位:
Ultra-Low-Power Complementary Logic with On-Chip Directly Assembled, Highly Adaptive 2-D Graphitic Platform
-
批准号:1002228
-
项目类别:Standard Grant
-
资助金额:$36.0万
-
财政年份:2010
-
负责人:Bin Yu
-
依托单位:
Collaborative Research: Multi-Level Behavior, Material Scalability and Energy Efficiency of 1-D Phase-Change Nanostructures
-
批准号:1005793
-
项目类别:Continuing Grant
-
资助金额:$29.0万
-
财政年份:2010
-
负责人:Bin Yu
-
依托单位:
NanoExcitonics: Implementing Basic Circuit Elements on 2-D Carbon System
-
批准号:1028267
-
项目类别:Standard Grant
-
资助金额:$36.0万
-
财政年份:2010
-
负责人:Bin Yu
-
依托单位:
Inference in high-dimension: statistics, computation and information theory
-
批准号:0907632
-
项目类别:Continuing Grant
-
资助金额:$20.0万
-
财政年份:2009
-
负责人:Bin Yu
-
依托单位:
CDI Type II: Collaborative Research: Sparse Inference: New Tools for Structural Knowledge Discovery
-
批准号:0835531
-
项目类别:Standard Grant
-
资助金额:$134.17万
-
财政年份:2008
-
负责人:Bin Yu
-
依托单位:
High-Dimensional Challenges in Statistical Machine Learning: Theory, Models and Algorithms
-
批准号:0605165
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2006
-
负责人:Bin Yu
-
依托单位:
Neural Coding in Visual and Auditory Systems for Natural Stimuli: Mathematical Modeling based on Data
-
批准号:0426227
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2006
-
负责人:Bin Yu
-
依托单位:
Boosting, Support Vector Machines, and Cloud Detection over Ice and Snow
-
批准号:0306508
-
项目类别:Standard Grant
-
资助金额:$27.5万
-
财政年份:2003
-
负责人:Bin Yu
-
依托单位:
Compressing and Analyzing Microarray Images for Genetic Information Extraction
-
批准号:0106656
-
项目类别:Continuing Grant
-
资助金额:$34.08万
-
财政年份:2001
-
负责人:Bin Yu
-
依托单位:
Discriminant Analysis of Hyperspectral Data for Bio Species Recognition
-
批准号:9803063
-
项目类别:Standard Grant
-
资助金额:$7.5万
-
财政年份:1998
-
负责人:Bin Yu
-
依托单位:
Adaptive Estimation in Wavelet Image Compression: Thresholding and Quantization
-
批准号:9802314
-
项目类别:Standard Grant
-
资助金额:$7.5万
-
财政年份:1998
-
负责人:Bin Yu
-
依托单位:
海外基金