课题基金 / 基金详情

Making Use of the Curse of Dimensionality in Modern Data Analysis

Making Use of the Curse of Dimensionality in Modern Data Analysis
在现代数据分析中利用维度诅咒
批准号:
2311399
负责人:
Hao Chen
金额:
$27.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31

项目摘要

项目成果

Hao Chen的其他基金

相似基金

相关文献

中文摘要
翻译
该研究项目深入研究尖端数据分析,经常处理高维或非欧几里德数据,如传感器读数,基因组信息,图像和网络数据集。 这些数据在不同学科中经常遇到,包括生物学,社会科学,计算机科学和天文学。 分析这些数据的一个主要挑战是维数灾难,这导致传统工具随着维度或特征数量的增加而迅速退化。 虽然以前的尝试集中在减少数据的维数或通过正则化技术,这些方法往往表现出局限性。 相比之下,这个项目采用了一种创新的策略,利用由于维度灾难而出现的模式来支持数据分析。 该项目旨在为数据分析提供有价值的工具,并探索统计在大数据时代的作用。 在该项目内开发的工具将作为开放源码软件包提供,并附有详尽的文件,加强统计界与各科学领域研究人员之间的合作,并使数据分析程序更加透明。 该项目还包括对本科生和研究生的培训和教育部分,使他们具备跨学科的数据分析技能,这对下一代研究人员来说是非常宝贵的。该项目旨在为涉及高维和非欧几里得数据的关键数据分析任务开发创新方法和基础理论。 具体而言,该项目将创建一个开创性的高维分类框架,该框架利用点间距离等级并考虑到维度灾难,从而与各种设置的现有方法相比,显着降低了错误分类率。 此外,该项目将建立一个统一的社区检测框架,能够识别所有三种社区结构- 通过将这些不同的结构链接到高维数据行为,其中由于维数灾难而自然出现核心-外围结构,统一框架在所有三种混合模式的模拟和真实的数据集上的数值研究中表现出上级性能,而现有方法在这些模式中的至少一种中挣扎。 最后,该项目将开发一个创新的高维聚类框架,采用维数灾难的模式,以减少误聚类率。 这些方法论和理论上的进步将增强对来自不同领域的现代复杂数据的理解,促进对这些领域中重大科学问题的理解。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This research project delves into cutting-edge data analysis, frequently dealing with high-dimensional or non-Euclidean data, such as sensor readings, genomic information, imagery, and network datasets. Such data is commonly encountered across diverse disciplines, including biology, social science, computer science, and astronomy. A major challenge in analyzing this data is the curse of dimensionality, which causes traditional tools to degrade rapidly as the number of dimensions or features grows. While previous attempts have focused on reducing the dimensionality of the data or through regularization techniques, these methods often exhibit limitations. In contrast, this project adopts an innovative strategy by harnessing the patterns that emerge as a result of the curse of dimensionality to bolster data analysis. The project aims to provide valuable tools for data analysis and explore the role of statistics in the era of big data. The tools developed within this project will be made available as open-source software packages with thorough documentation, enhancing collaboration between the statistics community and researchers from various scientific fields and making data analysis procedures more transparent. The project also includes training and educational components for undergraduate and graduate students, equipping them with interdisciplinary data analysis skills that will be invaluable for the next generation of researchers.This project seeks to develop innovative methodologies and foundational theories for crucial data analysis tasks involving high-dimensional and non-Euclidean data. Specifically, the project will create a pioneering high-dimensional classification framework that leverages interpoint distance ranks and takes into account the curse of dimensionality, resulting in significantly reduced misclassification rates compared to existing methods across various settings. Moreover, the project will establish a unified community detection framework capable of identifying all three community structures -- assortative, disassortative, and core-periphery – without prior knowledge of which community structure the network is. By linking these distinct structures to high-dimensional data behaviors, where the core-periphery structure naturally emerges due to the curse of dimensionality, the unified framework exhibits superior performance in numerical studies on simulated and real datasets across all three mixing patterns, whereas existing methods struggle in at least one of these patterns. Lastly, the project will develop an innovative high-dimensional clustering framework that employs the patterns of curse of dimensionality to reduce misclustering rates. These methodological and theoretical advancements will enhance the understanding of modern, complex data from diverse fields, promoting the comprehension of significant scientific issues in these areas.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ERI: Representations of Complex Engineering Systems via Technology Recursion and Renormalization Group
  • 批准号:
    2301627
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2023
  • 负责人:
    Hao Chen
  • 依托单位:
Development of Absolute Quantitative Protein Footprinting Mass Spectrometry (aqPFMS) for Probing Protein 3D Structures
  • 批准号:
    2203284
  • 项目类别:
    Standard Grant
  • 资助金额:
    $39.0万
  • 财政年份:
    2022
  • 负责人:
    Hao Chen
  • 依托单位:
SaTC: CORE: Small: Collaborative: Understanding and Detecting Memory Bugs in Rust
  • 批准号:
    1956364
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2020
  • 负责人:
    Hao Chen
  • 依托单位:
CAREER: New Change-Point Problems in Analyzing High-Dimensional and Non-Euclidean Data
  • 批准号:
    1848579
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2019
  • 负责人:
    Hao Chen
  • 依托单位:
海外基金