课题基金 / 基金详情

Collaborative Research: New Statistical Methods for Microbiome Data Analysis

Collaborative Research: New Statistical Methods for Microbiome Data Analysis
合作研究:微生物组数据分析的新统计方法
批准号:
2113360
负责人:
Jun Chen
金额:
$9.2万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-01 至 2024-08-31

项目摘要

项目成果

Jun Chen的其他基金

相似基金

相关文献

中文摘要
翻译
人类微生物群是与人体相关的微生物的集合,越来越多地被认为是人类健康和疾病的重要参与者。人类微生物组研究的重点是破译微生物组和宿主之间的复杂关系,并识别用于疾病预防、诊断和治疗的微生物生物标志物。目前研究人类微生物组的技术包括对样本中的微生物DNA进行测序,根据DNA测序可以确定微生物的身份和丰度。对这样的微生物组测序数据的分析提出了许多统计学挑战。首先,数据是零膨胀的。典型的微生物组数据集包含75%以上的零。其次,这些数据是由成分组成的。一种微生物的丰度变化会自动导致其他微生物的相对丰度的变化,从而使“驱动”微生物的识别变得困难。第三,这些微生物在系统发育上是相关的。密切相关的微生物通常具有相似的生物学特征。最后,人类的微生物群受到许多环境因素的影响。控制这些混杂因素对于做出有效的统计推断至关重要。该项目将开发新的统计方法来分析微生物组数据,以应对这些挑战。研究成果将通过科学出版物以及研讨会和会议报告加以传播。PI将通过GitHub和CRAN为已开发的方法开发、分发、记录和维护R软件包,并提供带有示例数据集的教程。PI将在真实世界的环境中彻底测试软件。鉴于研究人类微生物组的多重组学方法的流行,所提供的软件包将特别引起微生物组研究人员的兴趣。PIS将在高维统计、优化和基因组学的交叉点上培训学生。该项目有两个研究主题。在第一个推力中,PIS将为微生物组数据开发一个新的统计学习框架,以同时处理高维、成分效应、零膨胀和系统发育信息。特别是,新的框架包括一种基于新的Dirichlet混合模型的新的零补偿方法,一种在有监督/无监督统计学习中处理成分效应的通用方法,以及一种稳健的结构自适应方法来结合系统发生树中编码的外部信息。在第二个推力中,PI将开发一种二维错误发现率(FDR)控制程序,用于在微生物组关联分析中进行强大的混杂调节。该过程使用来自未调整分析的测试统计作为辅助统计,以过滤出大量不相关的特征,然后基于来自经调整分析的测试统计在约简集合上执行错误发现率控制。PIS将调查基于模型和无模型的方法,并证明渐近的FDR控制。这一裁决反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The human microbiome, the collection of micro-organisms associated with the human body, has been increasingly recognized as an important player in human health and disease. Human microbiome research focuses on deciphering the intricate relationship between the microbiome and the host and identifying microbial biomarkers for disease prevention, diagnosis, and treatment. Current technologies to study the human microbiome involve sequencing the microbial DNA in the sample, upon which the identity and the abundance of the micro-organisms can be determined. Analysis of such microbiome sequencing data raises many statistical challenges. First, the data are zero-inflated. A typical microbiome dataset contains more than 75% zeros. Second, the data are compositional. The abundance change in one microbe will automatically lead to changes in the relative abundance of others, making identification of the "driver" microbe difficult. Third, the microbes are phylogenetically related. Closely related microbes usually share similar biological traits. Finally, the human microbiome is subject to many environmental confounders. Controlling these confounders is essential to make valid statistical inferences. The project will develop novel statistical methods for analyzing microbiome data addressing these challenges. The research results will be disseminated through scientific publications as well as seminar and conference presentations. The PIs will develop, distribute, document, and maintain R software packages via GitHub and CRAN for developed methods, and provide tutorials with example datasets. The PIs will test the software in real-world settings thoroughly. Given the popularity of the multi-omics approach to study the human microbiome, the delivered software packages will be of particular interest to microbiome investigators. The PIs will train students at the intersection of high-dimensional statistics, optimization, and genomics.The project has two research thrusts. In the first thrust, the PIs will develop a new statistical learning framework for microbiome data to simultaneously tackle the high-dimensionality, compositional effect, zero-inflation, and phylogenetic information. In particular, the new framework includes a novel zero imputation method based on a new Dirichlet mixture model, a general approach for handling compositional effect in supervised/unsupervised statistical learning, and a robust structure adaptive method to incorporate external information encoded in the phylogenetic tree. In the second thrust, the PIs will develop a two-dimensional false discovery rate (FDR) control procedure for powerful confounder adjustment in microbiome association analysis. The procedure uses the test statistics from the unadjusted analysis as auxiliary statistics to filter out a large number of irrelevant features, and false discovery rate control is then performed based on the test statistics from the adjusted analysis on the reduced set. The PIs will investigate both model-based and model-free approaches, and prove the asymptotic FDR control.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
Benchmarking differential abundance analysis methods for correlated microbiome sequencing data
相关微生物组测序数据差异丰度分析方法的基准测试
DOI: 10.1093/bib/bbac607
发表时间: 2023
期刊: Briefings in Bioinformatics
影响因子: 9.5
作者: [Yang, Lu, Chen, Jun]
通讯作者: Chen, Jun
dICC: distance-based intraclass correlation coefficient for metagenomic reproducibility studies
dICC:用于宏基因组重现性研究的基于距离的组内相关系数
DOI: 10.1093/bioinformatics/btac618
发表时间: 2022
期刊: Bioinformatics
影响因子: 5.8
作者: [Chen, Jun, Zhang, Xianyang, Schwartz, ed., Russell]
通讯作者: Schwartz, ed., Russell
DOI: 10.1093/bioinformatics/btab498
发表时间: 2021-07-13
期刊: BIOINFORMATICS
影响因子: 5.8
作者: [Chen, Jun, Zhang, Xianyang]
通讯作者: Zhang, Xianyang
I-Corps: Wearable Magnetoelastic Generator for Atrial Fibrillation
CAREER: Reconfigurable and Predictive Control with Reinforcement Learning Supervisor for Active Battery Cell Balancing
  • 批准号:
    2237317
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2023
  • 负责人:
    Jun Chen
  • 依托单位:
ERI: Towards Safe Aviation Autonomy: A Risk-bounded Planning Framework for Dynamical Systems under Uncertainties
EAGER: Development of a Novel Rotating Wind Tunnel for 3D Study of Turbulent Flow
  • 批准号:
    2026329
  • 项目类别:
    Standard Grant
  • 资助金额:
    $25.92万
  • 财政年份:
    2020
  • 负责人:
    Jun Chen
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)