课题基金 / 基金详情

RI: Medium: Collaborative Research:Algorithmic High-Dimensional Statistics: Optimality, Computtional Barriers, and High-Dimensional Corrections

RI: Medium: Collaborative Research:Algorithmic High-Dimensional Statistics: Optimality, Computtional Barriers, and High-Dimensional Corrections
RI:中:协作研究:算法高维统计:最优性、计算障碍和高维校正
批准号:
2218713
负责人:
Yuxin Chen
金额:
$38.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-01-01 至 2024-07-31

项目摘要

项目成果

Yuxin Chen的其他基金

相似基金

相关文献

中文摘要
翻译
这项研究旨在解决从大维数据中学习和推理的紧迫挑战。当代传感和数据采集技术以前所未有的速度产生数据。因此,现代数据应用中普遍存在的一个挑战是如何高效、可靠地从海量数据中提取相关信息和相关见解。与此同时,人们需要推理的相关特征的前所未有的增长加剧了这一挑战,这往往甚至超过了数据样本的增长。对于机器学习和大数据分析的许多新兴应用来说,要么只在存在大量数据样本的情况下工作,要么完全忽略估计器的计算成本的经典统计推理范例,变得非常不足,甚至不可靠。为了在高维上解决上述紧迫问题,需要引入新的理论工具,以便全面了解各种算法和任务的性能限制。这个项目的目标有四个:首先,开发一种现代理论来表征经典统计算法在高维上的精确性能。其次,建议对经典的统计推断程序进行适当的修正,以适应样本匮乏的体制。第三,开发计算效率高的算法,如果可能的话,可以证明可以达到基本的统计极限。最后,第四,如果不能达到基本的统计限制,确定潜在的计算障碍。拟议研究计划的变革潜力在于通过统计学、近似理论、统计物理、数学优化和信息论的新组合来发展基本的统计数据分析理论,提供可扩展的统计推理和学习算法。在该项目中开发的理论和算法将对各种工程和科学应用产生直接影响,如大规模机器学习、DNA测序、遗传病分析和自然语言处理。这一合作项目为学生提供了跨大学的培训机会,我们致力于通过长期的导师和外展活动,吸引和帮助STEM中代表不足的学生和女性学生。本研究旨在解决从大维数据学习和推理方面的紧迫挑战。当代传感和数据采集技术以前所未有的速度产生数据。因此,现代数据应用中普遍存在的一个挑战是如何高效、可靠地从海量数据中提取相关信息和相关见解。与此同时,人们需要推理的相关特征的前所未有的增长加剧了这一挑战,这往往甚至超过了数据样本的增长。对于机器学习和大数据分析的许多新兴应用来说,要么只在存在大量数据样本的情况下工作,要么完全忽略估计器的计算成本的经典统计推理范例,变得非常不足,甚至不可靠。为了在高维上解决上述紧迫问题,需要引入新的理论工具,以便全面了解各种算法和任务的性能限制。这个项目的目标有四个:首先,开发一种现代理论来表征经典统计算法在高维上的精确性能。其次,建议对经典的统计推断程序进行适当的修正,以适应样本匮乏的体制。第三,开发计算效率高的算法,如果可能的话,可以证明可以达到基本的统计极限。最后,第四,如果不能达到基本的统计限制,确定潜在的计算障碍。拟议研究计划的变革潜力在于通过统计学、近似理论、统计物理、数学优化和信息论的新组合来发展基本的统计数据分析理论,提供可扩展的统计推理和学习算法。在该项目中开发的理论和算法将对各种工程和科学应用产生直接影响,如大规模机器学习、DNA测序、遗传病分析和自然语言处理。这一合作项目为学生提供了跨大学的培训机会,我们致力于通过长期的导师和外展活动,吸引和帮助STEM中代表不足的学生和女性学生。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This research aims to address the pressing challenges on learning and inference from large-dimensional data. Contemporary sensing and data acquisition technologies produce data at an unprecedented rate. A ubiquitous challenge in modern data applications is thus to efficiently and reliably extract relevant information and associated insights from a deluge of data. In the meantime, this challenge is exacerbated by the unprecedented growth of relevant features one needs to reason about, which oftentimes even outpaces the growth of data samples. Classical statistical inference paradigms, which either only work in the presence of an enormous number of data samples, or ignore the computational cost of the estimators at all, become highly insufficient, or even unreliable, for many emerging applications of machine learning and big-data analytics. To address the above pressing issues in high dimensions, novel theoretical tools need to be brought in the picture in order to provide a comprehensive understanding of the performance limits of various algorithms and tasks. The goal of this project is four-fold: First, to develop a modern theory to characterize precise performance of classical statistical algorithms in high dimensions. Second, to suggest proper corrections of classical statistical inference procedures to accommodate the sample-starved regime. Third, to develop computationally efficient algorithms that can provably attain the fundamental statistical limits, if possible. Finally, forth, to identify potential computational barriers if the fundamental statistical limits cannot be met. The transformative potential of the proposed research program is in the development of foundational statistical data analytics theory through a novel combination of statistics, approximation theory, statistical physics, mathematical optimization, and information theory, offering scalable statistical inference and learning algorithms. The theory and algorithms developed within this project will have direct impact on various engineering and science applications such as large-scale machine learning, DNA sequencing, genetic disease analysis, and natural language processing. This collaborative program provides cross-university opportunities for students training, and we are committed to engaging and helping underrepresented and women students in STEM through long-term mentorships and outreach activities.This research aims to address the pressing challenges on learning and inference from large-dimensional data. Contemporary sensing and data acquisition technologies produce data at an unprecedented rate. A ubiquitous challenge in modern data applications is thus to efficiently and reliably extract relevant information and associated insights from a deluge of data. In the meantime, this challenge is exacerbated by the unprecedented growth of relevant features one needs to reason about, which oftentimes even outpaces the growth of data samples. Classical statistical inference paradigms, which either only work in the presence of an enormous number of data samples, or ignore the computational cost of the estimators at all, become highly insufficient, or even unreliable, for many emerging applications of machine learning and big-data analytics. To address the above pressing issues in high dimensions, novel theoretical tools need to be brought in the picture in order to provide a comprehensive understanding of the performance limits of various algorithms and tasks. The goal of this project is four-fold: First, to develop a modern theory to characterize precise performance of classical statistical algorithms in high dimensions. Second, to suggest proper corrections of classical statistical inference procedures to accommodate the sample-starved regime. Third, to develop computationally efficient algorithms that can provably attain the fundamental statistical limits, if possible. Finally, forth, to identify potential computational barriers if the fundamental statistical limits cannot be met. The transformative potential of the proposed research program is in the development of foundational statistical data analytics theory through a novel combination of statistics, approximation theory, statistical physics, mathematical optimization, and information theory, offering scalable statistical inference and learning algorithms. The theory and algorithms developed within this project will have direct impact on various engineering and science applications such as large-scale machine learning, DNA sequencing, genetic disease analysis, and natural language processing. This collaborative program provides cross-university opportunities for students training, and we are committed to engaging and helping underrepresented and women students in STEM through long-term mentorships and outreach activities.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/tit.2022.3205781
发表时间: 2020-06
期刊: IEEE Transactions on Information Theory
影响因子: 2.5
作者: [Changxiao Cai;H. Poor;Yuxin Chen]
通讯作者: Changxiao Cai;H. Poor;Yuxin Chen
DOI: 10.1287/opre.2021.2106
发表时间: 2021-06-03
期刊: OPERATIONS RESEARCH
影响因子: 2.7
作者: [Cai, Changxiao, Li, Gen, Chen, Yuxin]
通讯作者: Chen, Yuxin
Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning
打破样本复杂性障碍,实现后悔最优无模型强化学习
DOI: 10.1093/imaiai/iaac034
发表时间: 2023
期刊: Information and Inference: A Journal of the IMA
影响因子: --
作者: [Li, Gen, Shi, Laixi, Chen, Yuxin, Chi, Yuejie]
通讯作者: Chi, Yuejie
DOI: 10.1109/tit.2021.3120096
发表时间: 2020-06
期刊: IEEE Transactions on Information Theory
影响因子: 2.5
作者: [Gen Li;Yuting Wei;Yuejie Chi;Yuantao Gu;Yuxin Chen]
通讯作者: Gen Li;Yuting Wei;Yuejie Chi;Yuantao Gu;Yuxin Chen
Collaborative Research: RI: Small: Foundations of Few-Round Active Learning
  • 批准号:
    2313131
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2023
  • 负责人:
    Yuxin Chen
  • 依托单位:
Collaborative Research: CIF: Medium: Statistical and Algorithmic Foundations of Efficient Reinforcement Learning
  • 批准号:
    2221009
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2022
  • 负责人:
    Yuxin Chen
  • 依托单位:
RI: Small: Uncertainty Quantification for Nonconvex Low-Complexity Models
  • 批准号:
    2218773
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2022
  • 负责人:
    Yuxin Chen
  • 依托单位:
Collaborative Research: CIF: Medium: Statistical and Algorithmic Foundations of Efficient Reinforcement Learning
  • 批准号:
    2106739
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2021
  • 负责人:
    Yuxin Chen
  • 依托单位:
海外基金