课题基金 / 基金详情

Methodology Development and Implementation for Microbiome Sequencing Data: Hierarchical Modeling on Clustered Taxa Counts with Repeated Measures

Methodology Development and Implementation for Microbiome Sequencing Data: Hierarchical Modeling on Clustered Taxa Counts with Repeated Measures
微生物组测序数据的方法开发和实施:重复测量的聚类分类群计数的分层建模
批准号:
RGPIN-2017-06672
负责人:
Xu, Wei
金额:
$2.04万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31

项目摘要

项目成果

Xu, Wei的其他基金

相似基金

相关文献

中文摘要
翻译
下一代测序的技术进步使研究人员能够揭示微生物群落的广泛变异性及其与不同疾病的关系。因此,了解影响微生物组组成的环境和宿主遗传因素变得至关重要。然而,由于微生物组测序数据的复杂性,该领域的强大和强大的方法还不发达,其中包括:a)微生物分类单元数据通常被分组为操作分类单元(OTU),并且这些计数通常是高度偏斜的、过度分散的和零膨胀的,B)分类层次聚类内的OTU计数通常是高度相关的,但是这种多变量性质通常被忽略,c)研究设计通常涉及从相关家庭成员中重复测量,从而引起时间和家庭相关性。在这个建议中,我将开发强大的生物信息学,统计和计算方法来克服这些挑战。具体而言,我建议使用潜变量(LV)的方法来共同建模多个分类群的纵向家庭研究框架内的层次分类集群。LV框架代表了集群的基本概念特征,并解释了不同类群之间的相关性。为了解决分类群计数的过度分散和零膨胀特征,我将在多变量OTU结果上应用零膨胀和障碍模型。 LV推断将基于贝叶斯框架构建,并从使用马尔可夫链蒙特卡罗(MCMC)算法获得的后验分布中进行采样。将开发贝叶斯模型选择算法,以选择特定数据集的最佳模型。我将结合降维方法的遗传因素,使遗传关联信号可以从全基因组数据中识别。我还将探索基因-基因(GxG)和基因-环境(GxE)在微生物组数据上的相互作用。高效的计算算法将使用C++开发,计算软件将在用户友好的界面中实现,并将分发给微生物组研究界。此外,还将构建用于微生物组数据建模和分析的标准化分析管道,并通过模拟进行测试。本研究亦将提供样本量估计及功效分析,以配合未来研究之设计。该提案将有助于规范和优化未来关于可改变的环境风险因素以及遗传因素的研究,以用于微生物组测序研究。该研究计划将为加拿大和国际遗传学和计算生物学研究人员社区推进大规模微生物组测序分析技术。
英文摘要
The technological advances in next generation sequencing have enabled researchers to unveil the wide variability in microbial communities and their relationships with different diseases. Therefore, it is becoming critical to understand both environmental and host genetic factors that impact the composition of the microbiome. However, robust and powerful methods in this area are underdeveloped due to the complexity of microbiome sequencing data, which includes: a) microbial taxa data are usually grouped into operational taxonomic units (OTUs) and these counts are often highly skewed, over-dispersed, and zero inflated, b) OTU counts within a taxonomic hierarchical cluster are often highly correlated, but this multivariate nature is usually ignored, c) the study designs often involve repeated measures taken from related family members, thus inducing temporal and familial correlations. In this proposal, I will develop powerful bioinformatics, statistical, and computational methods to overcome these challenges. Specifically, I propose to use the latent variable (LV) methodology to jointly model multiple taxa from hierarchical taxonomic clusters within a longitudinal family study framework. The LV framework represents the underlying conceptual traits of the cluster and explains the correlations among different taxa. To address the over-dispersed and zero inflated features of the taxa counts, I will apply both zero-inflated and hurdle models on the multivariate OTU outcomes. The LV inference will be constructed based on a Bayesian framework with samplings from the posterior distribution obtained using Markov Chain Monte Carlo (MCMC) algorithms. A Bayesian model selection algorithm will be developed to choose the optimal models for a particular dataset. I will incorporate dimensionality reduction methodologies on the genetic factors so that the genetic association signals can be identified from genome-wide data. I will also explore gene-gene (GxG), and gene-environment (GxE) interactions on the microbiome data. High-efficiency computational algorithms will be developed using C++, and computational software will be implemented within a user-friendly interface which will be distributed to the microbiome research community. In addition, a standardized analytic pipeline for modeling and analysis of microbiome data will be constructed and tested by simulations. Sample size estimation and power analysis based on both theoretical deduction and empirical results will also be provided to allow design of future studies. This proposal will help standardize and optimize future research on modifiable environmental risk factors, as well as genetic factors, for microbiome sequencing studies. This research program will advance large-scale microbiome sequencing analytic technologies for Canadian and international genetics and computational biology researcher community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Developing a Model-free Data-driven Framework for Problems in Finance
  • 批准号:
    RGPIN-2020-04686
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2022
  • 负责人:
    Xu, Wei
  • 依托单位:
Developing a Model-free Data-driven Framework for Problems in Finance
  • 批准号:
    RGPIN-2020-04686
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2021
  • 负责人:
    Xu, Wei
  • 依托单位:
Methodology Development and Implementation for Microbiome Sequencing Data: Hierarchical Modeling on Clustered Taxa Counts with Repeated Measures
  • 批准号:
    RGPIN-2017-06672
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $4.08万
  • 财政年份:
    2021
  • 负责人:
    Xu, Wei
  • 依托单位:
Developing a Model-free Data-driven Framework for Problems in Finance
  • 批准号:
    RGPIN-2020-04686
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2020
  • 负责人:
    Xu, Wei
  • 依托单位:
国内基金
海外基金
水稻边界发育缺陷突变体abnormal boundary development(abd)的基因克隆与功能分析
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位: