课题基金 / 基金详情

Making Use of the Curse of Dimensionality in Modern Data Analysis

Making Use of the Curse of Dimensionality in Modern Data Analysis
在现代数据分析中利用维度诅咒
批准号:
2311399
负责人:
Hao Chen
金额:
$27.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31

项目摘要

项目成果

Hao Chen的其他基金

相似基金

相关文献

中文摘要
翻译
该研究项目深入研究前沿数据分析,经常处理高维或非欧几里得数据,如传感器读数、基因组信息、图像和网络数据集。这样的数据在不同的学科中都很常见,包括生物学、社会科学、计算机科学和天文学。分析这些数据的一个主要挑战是维度诅咒,它导致传统工具随着维度或特征的数量增长而迅速降级。虽然以前的尝试侧重于降低数据的维度或通过正则化技术,但这些方法往往显示出局限性。相比之下,这个项目采用了一种创新的战略,利用因维度诅咒而出现的模式来支持数据分析。该项目旨在为数据分析提供有价值的工具,并探索统计在大数据时代的作用。在该项目内开发的工具将以开放源码软件包的形式提供,并提供详尽的文件,从而加强统计界与不同科学领域的研究人员之间的合作,并使数据分析程序更加透明。该项目还包括本科生和研究生的培训和教育部分,使他们掌握跨学科的数据分析技能,这对下一代研究人员将是无价的。该项目寻求为涉及高维和非欧几里德数据的关键数据分析任务开发创新的方法和基础理论。具体地说,该项目将创建一个开创性的高维分类框架,该框架利用点间距离等级并考虑到维度诅咒,与各种环境下的现有方法相比,大大降低了错误分类率。此外,该项目将建立一个统一的社区检测框架,能够在事先不知道网络是哪种社区结构的情况下识别所有三种社区结构--分类、非分类和核心-边缘。通过将这些不同的结构与高维数据行为联系起来,其中核心-外围结构由于维度诅咒而自然出现,统一框架在对所有三种混合模式的模拟和真实数据集的数值研究中显示出优越的性能,而现有方法至少在这些模式中的一种模式中挣扎。最后,该项目将开发一个创新的高维集群框架,采用维度灾难的模式来减少错群率。这些方法和理论的进步将加强对来自不同领域的现代、复杂数据的理解,促进对这些领域重大科学问题的理解。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This research project delves into cutting-edge data analysis, frequently dealing with high-dimensional or non-Euclidean data, such as sensor readings, genomic information, imagery, and network datasets. Such data is commonly encountered across diverse disciplines, including biology, social science, computer science, and astronomy. A major challenge in analyzing this data is the curse of dimensionality, which causes traditional tools to degrade rapidly as the number of dimensions or features grows. While previous attempts have focused on reducing the dimensionality of the data or through regularization techniques, these methods often exhibit limitations. In contrast, this project adopts an innovative strategy by harnessing the patterns that emerge as a result of the curse of dimensionality to bolster data analysis. The project aims to provide valuable tools for data analysis and explore the role of statistics in the era of big data. The tools developed within this project will be made available as open-source software packages with thorough documentation, enhancing collaboration between the statistics community and researchers from various scientific fields and making data analysis procedures more transparent. The project also includes training and educational components for undergraduate and graduate students, equipping them with interdisciplinary data analysis skills that will be invaluable for the next generation of researchers.This project seeks to develop innovative methodologies and foundational theories for crucial data analysis tasks involving high-dimensional and non-Euclidean data. Specifically, the project will create a pioneering high-dimensional classification framework that leverages interpoint distance ranks and takes into account the curse of dimensionality, resulting in significantly reduced misclassification rates compared to existing methods across various settings. Moreover, the project will establish a unified community detection framework capable of identifying all three community structures -- assortative, disassortative, and core-periphery – without prior knowledge of which community structure the network is. By linking these distinct structures to high-dimensional data behaviors, where the core-periphery structure naturally emerges due to the curse of dimensionality, the unified framework exhibits superior performance in numerical studies on simulated and real datasets across all three mixing patterns, whereas existing methods struggle in at least one of these patterns. Lastly, the project will develop an innovative high-dimensional clustering framework that employs the patterns of curse of dimensionality to reduce misclustering rates. These methodological and theoretical advancements will enhance the understanding of modern, complex data from diverse fields, promoting the comprehension of significant scientific issues in these areas.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ERI: Representations of Complex Engineering Systems via Technology Recursion and Renormalization Group
  • 批准号:
    2301627
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2023
  • 负责人:
    Hao Chen
  • 依托单位:
Development of Absolute Quantitative Protein Footprinting Mass Spectrometry (aqPFMS) for Probing Protein 3D Structures
  • 批准号:
    2203284
  • 项目类别:
    Standard Grant
  • 资助金额:
    $39.0万
  • 财政年份:
    2022
  • 负责人:
    Hao Chen
  • 依托单位:
SaTC: CORE: Small: Collaborative: Understanding and Detecting Memory Bugs in Rust
  • 批准号:
    1956364
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2020
  • 负责人:
    Hao Chen
  • 依托单位:
CAREER: New Change-Point Problems in Analyzing High-Dimensional and Non-Euclidean Data
  • 批准号:
    1848579
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2019
  • 负责人:
    Hao Chen
  • 依托单位:
海外基金