Statistical and Computational Tools for Analyzing High-Dimensional Heterogeneous Data
Statistical and Computational Tools for Analyzing High-Dimensional Heterogeneous Data
批准号:
2210907
负责人:
Kaizheng Wang
金额:
$18.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-01 至 2025-07-31
中文摘要
现代技术以各种形式产生大量数据。高吞吐量的数据不可避免地伴随着巨大的异构性和大量的噪声。例如,大规模的遗传研究通常涉及具有各种不同属性的人;社交网络通常由多个隐藏的社区组成,与外部联系相比,内部联系更加紧密。虽然原始特征具有高环境维度(例如,数千个基因),但通常内在结构表现出低复杂性(例如,个体地理位置的纬度和经度)。对潜在结构的精确提取为解决下游任务铺平了道路。面对统计和计算方面的重大挑战,本项目旨在开发有效的方法,从异构数据中估计和推断潜在结构。该项目将产生用于科学研究的尖端工具,易于实施的开源软件,以及用于理论分析的新数学定理。该项目还将为统计教育和研究培训提供大量机会。在第一部分中,目标是开发一种新的灵活的方法来聚类高维数据。这一部分的目的是新的算法,可以识别非球形,甚至非凸集群。对混合模型的深入分析带来了理论见解,包括严格的有限样本统计误差界和有限迭代收敛保证计算。在第二部分中,目标是研究异构关系数据,这些数据将单个对象的信息编码在它们的成对关系中。这一部分产生了可靠的方法,估计和测试潜在的结构在现实的情况下,部分观察到的数据可能不会均匀随机抽样。最后,第三部分着重于多个相关数据集的联合分析,如具有高维个人属性的社交网络。该项目前两部分开发的工具将构成实现最后一个重点研究目标的基本构件。研究结果将为提高统计准确性提供新的有效的数据集成策略。拟议的研究举措包括以公开软件的形式传播新方法和算法,以及加强跨学科研究培训和加强统计科学多样性的积极议程,该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Modern technologies generate tremendous volumes of data in diverse forms. The high throughput data come inevitably with great heterogeneity and enormous amount of noise. For instance, a large-scale genetic study typically involves people with various different attributes; a social network usually consists of multiple hidden communities with denser internal connections compared to external ones. While the raw features have high ambient dimension (for example, thousands of genes), oftentimes the intrinsic structures exhibit low complexity (for example, latitude and longitude of an individual’s geographic location). Precise extraction of the latent structure paves the way for solving downstream tasks. Faced with the significant challenges in statistics and computation, this project aims to develop efficient methodologies for estimating and inferring latent structures from heterogeneous data. This project will yield cutting-edge tools for scientific study, open-source software for easy implementation, and new mathematical theorems for theoretical analysis. The project will also provide numerous opportunities for statistical education and research training.The project is structured into three parts. In the first part, the goal is to develop a new flexible methodology for clustering high-dimensional data. This part aims at new algorithms that can identify non-spherical and even non-convex clusters. An in-depth analysis of mixture models brings theoretical insights including tight finite-sample statistical error bounds and finite-iteration convergence guarantees for computation. In the second part, the goal is to study heterogeneous relational data that encode the information of individual objects in their pairwise relations. This part yields reliable methods for estimating and testing latent structures in the realistic scenario where the partially observed data may not be uniformly sampled at random. Finally, the third part focuses on the joint analysis of multiple related datasets, such as social networks with high-dimensional personal attributes. The tools developed in the first two parts of the project will constitute fundamental building blocks to address the research goals of this last thrust. The research finding will provide novel efficient data integration strategies for enhanced statistical accuracy. The proposed research initiatives include dissemination of the new methods and algorithms in a form of publicly available software and an active agenda on enhancing interdisciplinary research training and enhancing diversity in statistical sciences.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: