Statistical Methods for Analyzing Complex Structured and Count Data
Statistical Methods for Analyzing Complex Structured and Count Data
批准号:
2210019
负责人:
Fang Han
金额:
$20.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-15 至 2025-08-31
中文摘要
如何衡量复杂社会环境中的治疗/干预效果?单细胞数据如何用于诊断自闭症谱系障碍等复杂疾病?该项目旨在通过开发统计和机器学习方法来解决这些问题,这些方法能够对生物学、神经科学、社会科学、政治学和流行病学中常规生成的大数据集进行强大、可解释、高效和快速的分析。该项目包括两个主要方向:(1)大型数据集的结构化分析,其目标是设计出能够有效利用可能非常高维数据的内在结构而无需首先估计结构的方法;以及(2)大计数数据集的分析,其目标是设计鲁棒的非参数模型和算法,所述非参数模型和算法可以处理复杂的,可能是异构的,计数数据。调查员还计划指导和支持统计及相关领域的研究生和本科生,扩大代表性不足的少数民族学生的参与,该项目将通过提出两个主要研究方向,推进大规模结构化和计数数据分析方面的知识现状。第一种是通过最近邻(NN)或最小生成树进行基于随机图的统计推断。这条赛道中的两个主要工作示例是用于推断平均治疗效果的NN匹配和用于推断边际和条件依赖强度的基于图形的相关系数。研究人员的目标是修改和推广这两类方法,以提高它们的效率,同时保持它们的鲁棒性和计算速度。第二个是集中在非参数单变量或多变量泊松混合模型。研究者的目标是将异质计数值混合模型与非参数模型(例如,完全非参数、形状约束、基于非负矩阵分解等)基于模型的非均匀混合推理。研究者将在这两条轨道上探索和解决一些理论、方法、计算和应用问题。在第一个轨道上取得的一些初步成果已经激发了因果推理界的新工作,从第二个轨道产生的结果预计将有助于自闭症的早期诊断。这个奖项反映了NSF的法定使命,并已被认为是值得支持的评估使用基金会的智力价值和更广泛的影响审查标准。
英文摘要
How can treatment/intervention effects in a complex social environment be measured? How can single-cell data be used to diagnose complex diseases such as autism spectrum disorder? This project aims to address these questions by developing statistics and machine learning methods that enable robust, interpretable, efficient, and fast analysis of big datasets routinely produced in biology, neuroscience, social sciences, politics, and epidemiology. The project encompasses two main tracks: (1) structured analysis for large datasets, for which the goal is to devise methods that can make efficient use of the intrinsic structure of the possibly very high-dimensional data without having to estimate the structure first; and (2) analysis of large count datasets, for which the goal is to design robust nonparametric models and algorithms that can handle complex, likely heterogeneous, count data. The investigator also plans to mentor and support graduate and undergraduate students majoring in statistics and related fields and broaden the participation of underrepresented minority students.This project will advance the current state of knowledge in big structured and count data analyses by putting forward two main tracks of studies. The first is centered on random graph-based statistical inference through nearest neighbors (NN) or minimum spanning tree. Two main working examples in this track are NN matching for inferring the average treatment effect and graph-based correlation coefficients to infer marginal and conditional dependence strength. The investigator aims to revise and generalize these two families of methods to boost their efficiency while maintaining their robustness and computational speed. The second is centered on nonparametric univariate or multivariate Poisson mixture models. The investigator aims to bridge heterogeneous count-valued mixtures to nonparametric models (e.g., fully nonparametric, shape-constrained, nonnegative matrix factorization-based, etc.) under the umbrella of heterogeneous mixture model-based inference. The investigator will explore and settle several theory, method, computation, and application questions in the two tracks. Some preliminary results made in the first track have already stimulated new work in the causal inference community, and the results produced from the second track are expected to help with the early diagnosis of autism.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Rank-based Inference for Complex and Noisy High-dimensional Data
-
批准号:2019363
-
项目类别:Standard Grant
-
资助金额:$29.0万
-
财政年份:2020
-
负责人:Fang Han
-
依托单位:
An Integrated Toolkit for High-Dimensional Complex and Time Series Data Analysis
-
批准号:1712536
-
项目类别:Standard Grant
-
资助金额:$16.0万
-
财政年份:2017
-
负责人:Fang Han
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: