课题基金 / 基金详情

Collaborative Research: Towards designing optimal learning procedures via precise medium-dimensional asymptotic analysis

Collaborative Research: Towards designing optimal learning procedures via precise medium-dimensional asymptotic analysis
协作研究:通过精确的中维渐近分析设计最佳学习程序
批准号:
2210505
负责人:
Haolei Weng
金额:
$12.48万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-01 至 2025-07-31

项目摘要

项目成果

Haolei Weng的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的十年中,数据科学和人工智能(AI)已经成功地解决了医疗保健、教育和自主系统等各个生活领域面临的一些最重要的科学和工程挑战。最近的一个例子是谷歌DeepMind开发的深度学习程序AlphaFold,它可以从氨基酸序列中预测蛋白质的3D结构,其准确性与实验相当。尽管近年来取得了显着的进展,但大数据的有效统计学习程序的设计-现代数据科学和人工智能的核心组成部分-仍然是特设的,并且对这种设计的学习方案的精确理论理解处于起步阶段。特别是,以下基本问题仍然是开放的:(i)如何获得可操作的见解的学习算法的性能,而无需计算要求的实验方法?(ii)在给定的数据密集型环境中,什么是最佳学习过程?解决这些问题将为下一代数据科学和人工智能的发展铺平道路,最终有助于提高生活质量。该项目旨在开发一种新的分析方法,可以解决上述挑战。新的框架预计将建立定量的不同的学习算法的性能的精确表征,并提供了一个一般的配方设计最佳的学习proceedings.Most的国家的最先进的学习系统考虑复杂的模型,其中的参数,p的数量,是相当大的。在大多数情况下,p要么比使用中的数据中的观测数n大得多,要么与其相当。这种新的常规挑战了我们对科学和技术中无处不在的统计模型和程序的理论理解。一方面,基于n很大且p比n小得多的假设的经典分析技术不能为上述当代场景中的统计学习提供有效的预测。另一方面,现代的非渐近分析框架,已经非常成功的顺序明智的风险特征,往往达不到提供尖锐的结果。因此,对于各种统计模型,如何以最佳方式解决各种学习问题仍然很不清楚。该项目旨在填补这一空白,通过提供一个大家庭的统计模型,包括广义线性模型作为一个子集的精确的理论理解。它专注于中维区域,其中p与n成线性关系,并为研究不同学习方案的准确性和表征最佳性能创造了新的工具。该项目的预期成果是:(i)发现广泛的一类学习方法的精确性能限制和(ii)评估信息理论下限和现有算法的性能之间的差距。这些结果最终将为最佳学习程序的设计提供启示。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In the past decade, data science and artificial intelligence (AI) have successfully addressed some of the most important scientific and engineering challenges faced in various domains of life like healthcare, education and autonomous systems. A recent example is the deep learning program AlphaFold developed by Google’s DeepMind which can predict a protein’s 3D structure from its amino acid sequence with accuracy competitive to experiment. Despite remarkable progress made in recent years, the design of efficient statistical learning procedures for big data – a core component in modern data science and AI, has remained ad-hoc, and the precise theoretical understanding of such designed learning schemes is in its infancy. In particular, the following fundamental questions have remained open: (i) how to gain actionable insights into the performance of a learning algorithm without computationally demanding experimental methods? (ii) what is the optimal learning procedure in a given data-intensive environment? Answering such questions will pave the way for the development of the next generation of data science and AI, ultimately contributing to a better quality of life. This project aims to develop a novel analysis approach that can address the above challenges. The new framework is expected to establish quantitatively precise characterizations of the performance of diverse learning algorithms and provide a general recipe for designing optimal learning procedures.Most of the state-of-the-art learning systems consider sophisticated models in which the number of parameters, p, is substantially large. In most cases p is either much larger than or comparable to the number of observations, n, in the data in use. This new routine has challenged our theoretical understanding of ubiquitous statistical models and procedures in science and technology. On one hand, classical analysis techniques based on the assumption that n is large and p is much smaller than n do not provide valid predictions for statistical learning in the aforementioned contemporary scenarios. On the other hand, modern non-asymptotic analysis frameworks which have been very successful in order-wise risk characterizations, often fall short of delivering sharp results. Hence, it remains largely unclear how to solve various learning problems in an optimal fashion for a variety of statistical models. This project aims to fill this gap by providing a precise theoretical understanding of a large family of statistical models including generalized linear models as a subset. It focuses on the medium-dimensional regime where p scales linearly with n, and creates new tools for studying the accuracy of different learning schemes and characterizing the optimal performances. The expected outcomes of this project are: (i) discovering the precise performance limits of a broad class of learning methods and (ii) evaluating the gaps between information-theoretic lower bounds and performance of the existing algorithms. Such results will ultimately shed light on the design of optimal learning procedures. The proposal will also provide numerous opportunities for interdisciplinary research training and professional career development of future generation of statisticians.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Impact of Molecular Biomarkers on Survival Outcomes: Signal Detection, Model Selection, and Statistical Inference in High-dimensional Settings
  • 批准号:
    1915099
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $12.0万
  • 财政年份:
    2019
  • 负责人:
    Haolei Weng
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)