课题基金 / 基金详情

Collaborative Research: Towards designing optimal learning procedures via precise medium-dimensional asymptotic analysis

Collaborative Research: Towards designing optimal learning procedures via precise medium-dimensional asymptotic analysis
协作研究:通过精确的中维渐近分析设计最佳学习程序
批准号:
2210506
负责人:
Mohammad Ali Maleki
金额:
$13.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-01 至 2025-07-31

项目摘要

项目成果

Mohammad Ali Maleki的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的十年里,数据科学和人工智能(AI)成功地解决了医疗保健、教育和自主系统等生活各个领域面临的一些最重要的科学和工程挑战。最近的一个例子是谷歌DeepMind开发的深度学习程序AlphaFold,它可以根据蛋白质的氨基酸序列预测蛋白质的三维结构,其精度与实验相当。尽管近年来取得了显著进展,但为大数据设计高效的统计学习程序--现代数据科学和人工智能的核心组件--仍然是临时的,对这种设计的学习方案的准确理论理解还处于初级阶段。特别是,以下基本问题仍然悬而未决:(I)如何在没有计算要求的实验方法的情况下,对学习算法的性能获得可操作的见解?(2)在给定的数据密集型环境中,最佳学习程序是什么?回答这些问题将为下一代数据科学和人工智能的发展铺平道路,最终为更好的生活质量做出贡献。该项目旨在开发一种新的分析方法,以应对上述挑战。新的框架有望建立不同学习算法性能的定量精确表征,并提供设计最优学习过程的一般配方。大多数最先进的学习系统考虑其中参数p的数量相当大的复杂模型。在大多数情况下,p要么比正在使用的数据中的观测次数n大得多,要么与之相当。这一新的惯例挑战了我们对科学技术中无处不在的统计模型和程序的理论理解。一方面,基于n大而p远小于n的假设的经典分析技术不能为上述当代情景中的统计学习提供有效的预测。另一方面,现代非渐近分析框架在按顺序描述风险方面非常成功,但往往无法提供准确的结果。因此,对于各种统计模型,如何以最佳方式解决各种学习问题在很大程度上仍然不清楚。该项目旨在通过提供对一大类统计模型的精确理论理解来填补这一空白,这些模型包括作为子集的广义线性模型。它聚焦于p与n成线性关系的中维区域,并为研究不同学习方案的精度和表征最优性能创造了新的工具。该项目的预期结果是:(I)发现一大类学习方法的精确性能限制,以及(Ii)评估信息理论下限与现有算法性能之间的差距。这样的结果最终将有助于优化学习程序的设计。该提案还将为未来一代统计学家的跨学科研究培训和职业发展提供大量机会。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In the past decade, data science and artificial intelligence (AI) have successfully addressed some of the most important scientific and engineering challenges faced in various domains of life like healthcare, education and autonomous systems. A recent example is the deep learning program AlphaFold developed by Google’s DeepMind which can predict a protein’s 3D structure from its amino acid sequence with accuracy competitive to experiment. Despite remarkable progress made in recent years, the design of efficient statistical learning procedures for big data – a core component in modern data science and AI, has remained ad-hoc, and the precise theoretical understanding of such designed learning schemes is in its infancy. In particular, the following fundamental questions have remained open: (i) how to gain actionable insights into the performance of a learning algorithm without computationally demanding experimental methods? (ii) what is the optimal learning procedure in a given data-intensive environment? Answering such questions will pave the way for the development of the next generation of data science and AI, ultimately contributing to a better quality of life. This project aims to develop a novel analysis approach that can address the above challenges. The new framework is expected to establish quantitatively precise characterizations of the performance of diverse learning algorithms and provide a general recipe for designing optimal learning procedures.Most of the state-of-the-art learning systems consider sophisticated models in which the number of parameters, p, is substantially large. In most cases p is either much larger than or comparable to the number of observations, n, in the data in use. This new routine has challenged our theoretical understanding of ubiquitous statistical models and procedures in science and technology. On one hand, classical analysis techniques based on the assumption that n is large and p is much smaller than n do not provide valid predictions for statistical learning in the aforementioned contemporary scenarios. On the other hand, modern non-asymptotic analysis frameworks which have been very successful in order-wise risk characterizations, often fall short of delivering sharp results. Hence, it remains largely unclear how to solve various learning problems in an optimal fashion for a variety of statistical models. This project aims to fill this gap by providing a precise theoretical understanding of a large family of statistical models including generalized linear models as a subset. It focuses on the medium-dimensional regime where p scales linearly with n, and creates new tools for studying the accuracy of different learning schemes and characterizing the optimal performances. The expected outcomes of this project are: (i) discovering the precise performance limits of a broad class of learning methods and (ii) evaluating the gaps between information-theoretic lower bounds and performance of the existing algorithms. Such results will ultimately shed light on the design of optimal learning procedures. The proposal will also provide numerous opportunities for interdisciplinary research training and professional career development of future generation of statisticians.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Consistent Risk Estimation under High-Dimensional Asymptotics
  • 批准号:
    1810888
  • 项目类别:
    Standard Grant
  • 资助金额:
    $11.99万
  • 财政年份:
    2018
  • 负责人:
    Mohammad Ali Maleki
  • 依托单位:
CIF: Small: Collaborative Research: Towards universal signal recovery algorithms
  • 批准号:
    1420328
  • 项目类别:
    Standard Grant
  • 资助金额:
    $24.89万
  • 财政年份:
    2014
  • 负责人:
    Mohammad Ali Maleki
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)