Statistical Methods for Analyzing Complex Structured and Count Data
Statistical Methods for Analyzing Complex Structured and Count Data
批准号:
2210019
负责人:
Fang Han
金额:
$20.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-15 至 2025-08-31
中文摘要
如何衡量复杂社会环境中的治疗/干预效果?单细胞数据如何用于诊断复杂的疾病,如自闭症谱系障碍?该项目旨在通过开发统计和机器学习方法来解决这些问题,这些方法能够对生物学、神经科学、社会科学、政治学和流行病学中常规产生的大数据集进行健壮、可解释、高效和快速的分析。该项目包括两个主要方面:(1)对大型数据集的结构化分析,其目标是设计出能够有效利用可能非常高维数据的内在结构的方法,而不必首先估计结构;(2)对大型计数数据集的分析,其目标是设计能够处理复杂的、可能是异质的计数数据的健壮的非参数模型和算法。调查人员还计划指导和支持统计学及相关领域的研究生和本科生,并扩大未被充分代表的少数民族学生的参与。该项目将通过提出两个主要研究轨道来促进大型结构化和计数数据分析方面的知识现状。第一种是基于最近邻(NN)或最小生成树的基于随机图的统计推理。这一跟踪中的两个主要工作实例是用于推断平均治疗效果的神经网络匹配和用于推断边缘和条件性依赖强度的基于图形的相关系数。研究人员的目标是修改和推广这两类方法,以提高它们的效率,同时保持它们的健壮性和计算速度。第二种是以非参数单变量或多变量Poisson混合模型为中心的。研究人员的目标是将不同种类的计数值混合模型与非参数模型(例如,完全非参数、形状约束、基于非负矩阵分解等)联系起来。在异质混合的保护伞下,基于模型的推理。研究人员将在这两个轨道上探索和解决几个理论、方法、计算和应用问题。第一个赛道取得的一些初步结果已经刺激了因果推理界的新工作,第二个赛道产生的结果预计将有助于自闭症的早期诊断。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
How can treatment/intervention effects in a complex social environment be measured? How can single-cell data be used to diagnose complex diseases such as autism spectrum disorder? This project aims to address these questions by developing statistics and machine learning methods that enable robust, interpretable, efficient, and fast analysis of big datasets routinely produced in biology, neuroscience, social sciences, politics, and epidemiology. The project encompasses two main tracks: (1) structured analysis for large datasets, for which the goal is to devise methods that can make efficient use of the intrinsic structure of the possibly very high-dimensional data without having to estimate the structure first; and (2) analysis of large count datasets, for which the goal is to design robust nonparametric models and algorithms that can handle complex, likely heterogeneous, count data. The investigator also plans to mentor and support graduate and undergraduate students majoring in statistics and related fields and broaden the participation of underrepresented minority students.This project will advance the current state of knowledge in big structured and count data analyses by putting forward two main tracks of studies. The first is centered on random graph-based statistical inference through nearest neighbors (NN) or minimum spanning tree. Two main working examples in this track are NN matching for inferring the average treatment effect and graph-based correlation coefficients to infer marginal and conditional dependence strength. The investigator aims to revise and generalize these two families of methods to boost their efficiency while maintaining their robustness and computational speed. The second is centered on nonparametric univariate or multivariate Poisson mixture models. The investigator aims to bridge heterogeneous count-valued mixtures to nonparametric models (e.g., fully nonparametric, shape-constrained, nonnegative matrix factorization-based, etc.) under the umbrella of heterogeneous mixture model-based inference. The investigator will explore and settle several theory, method, computation, and application questions in the two tracks. Some preliminary results made in the first track have already stimulated new work in the causal inference community, and the results produced from the second track are expected to help with the early diagnosis of autism.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Rank-based Inference for Complex and Noisy High-dimensional Data
-
批准号:2019363
-
项目类别:Standard Grant
-
资助金额:$29.0万
-
财政年份:2020
-
负责人:Fang Han
-
依托单位:
An Integrated Toolkit for High-Dimensional Complex and Time Series Data Analysis
-
批准号:1712536
-
项目类别:Standard Grant
-
资助金额:$16.0万
-
财政年份:2017
-
负责人:Fang Han
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: