课题基金 / 基金详情

Statistical Methods for Analyzing Complex Structured and Count Data

Statistical Methods for Analyzing Complex Structured and Count Data
分析复杂结构化和计数数据的统计方法
批准号:
2210019
负责人:
Fang Han
金额:
$20.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-15 至 2025-08-31

项目摘要

项目成果

Fang Han的其他基金

相似基金

相关文献

中文摘要
翻译
如何在复杂的社会环境中衡量治疗/干预的效果?如何使用单细胞数据来诊断复杂的疾病,如自闭症谱系障碍?该项目旨在通过开发统计学和机器学习方法来解决这些问题,这些方法可以对生物学、神经科学、社会科学、政治学和流行病学中常规产生的大数据集进行稳健、可解释、高效和快速的分析。该项目包括两个主要方向:(1)大型数据集的结构化分析,其目标是设计出可以有效利用可能非常高维数据的内在结构的方法,而不必首先估计结构;(2)大型计数数据集的分析,其目标是设计稳健的非参数模型和算法,可以处理复杂的,可能是异构的计数数据。研究者还计划指导和支持统计及相关领域的研究生和本科生,并扩大代表性不足的少数民族学生的参与。本项目将通过提出两个主要的研究方向来推进大结构化和计数数据分析的知识现状。第一种是通过最近邻(NN)或最小生成树进行基于随机图的统计推断。这条轨道上的两个主要工作示例是用于推断平均处理效果的神经网络匹配和用于推断边缘和条件依赖强度的基于图的相关系数。研究者的目的是修改和推广这两个家族的方法,以提高他们的效率,同时保持他们的鲁棒性和计算速度。第二是集中在非参数的单变量或多变量泊松混合模型。研究者的目标是在基于异构混合模型的推理的保护伞下,将异构计数混合连接到非参数模型(例如,完全非参数,形状约束,基于非负矩阵分解等)。研究者将在两个轨道上探索和解决一些理论、方法、计算和应用问题。在第一个轨道上取得的一些初步结果已经刺激了因果推理界的新工作,而在第二个轨道上产生的结果有望有助于自闭症的早期诊断。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
How can treatment/intervention effects in a complex social environment be measured? How can single-cell data be used to diagnose complex diseases such as autism spectrum disorder? This project aims to address these questions by developing statistics and machine learning methods that enable robust, interpretable, efficient, and fast analysis of big datasets routinely produced in biology, neuroscience, social sciences, politics, and epidemiology. The project encompasses two main tracks: (1) structured analysis for large datasets, for which the goal is to devise methods that can make efficient use of the intrinsic structure of the possibly very high-dimensional data without having to estimate the structure first; and (2) analysis of large count datasets, for which the goal is to design robust nonparametric models and algorithms that can handle complex, likely heterogeneous, count data. The investigator also plans to mentor and support graduate and undergraduate students majoring in statistics and related fields and broaden the participation of underrepresented minority students.This project will advance the current state of knowledge in big structured and count data analyses by putting forward two main tracks of studies. The first is centered on random graph-based statistical inference through nearest neighbors (NN) or minimum spanning tree. Two main working examples in this track are NN matching for inferring the average treatment effect and graph-based correlation coefficients to infer marginal and conditional dependence strength. The investigator aims to revise and generalize these two families of methods to boost their efficiency while maintaining their robustness and computational speed. The second is centered on nonparametric univariate or multivariate Poisson mixture models. The investigator aims to bridge heterogeneous count-valued mixtures to nonparametric models (e.g., fully nonparametric, shape-constrained, nonnegative matrix factorization-based, etc.) under the umbrella of heterogeneous mixture model-based inference. The investigator will explore and settle several theory, method, computation, and application questions in the two tracks. Some preliminary results made in the first track have already stimulated new work in the causal inference community, and the results produced from the second track are expected to help with the early diagnosis of autism.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Rank-based Inference for Complex and Noisy High-dimensional Data
  • 批准号:
    2019363
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.0万
  • 财政年份:
    2020
  • 负责人:
    Fang Han
  • 依托单位:
An Integrated Toolkit for High-Dimensional Complex and Time Series Data Analysis
  • 批准号:
    1712536
  • 项目类别:
    Standard Grant
  • 资助金额:
    $16.0万
  • 财政年份:
    2017
  • 负责人:
    Fang Han
  • 依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data