课题基金 / 基金详情

Measures of association in microbiome and other count-based platforms

Measures of association in microbiome and other count-based platforms
微生物组和其他基于计数的平台的关联测量
批准号:
RGPIN-2021-03634
负责人:
McGregor, Kevin
金额:
$1.68万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

McGregor, Kevin的其他基金

相似基金

相关文献

中文摘要
翻译
人体微生物组是居住在人体内或体表的生物体的集合,在人体健康中起着不可分割的作用。典型的微生物组数据集由计数矩阵组成,其中行对应于来自不同微生物种群的样本,列对应于微生物分类群(例如种,属,科)。微生物组测序数据本质上是组成的,这意味着样品中一个分类群的计数不应该以绝对的术语来解释;相反,它应该只相对于一个或多个其他分类群进行解释。这使得测量不同分类群之间的关联变得困难,因为原始分类群比例的相关性将是负偏倚的。在这个建议中,我将扩展现有的多项模型来估计组合数据中的协方差矩阵作为协变量的函数。为此,我将修改之前开发的模型,以处理更灵活的协方差结构。与此相关的是,统计技术已经发展到解决组合数据的缺陷,例如对数比变换。然而,对数比变换是尺度不变的,而基于计数的测序平台则不是。最近,(Lovell, Chua, and McGrath 2020)发表了一篇重要的论文,探讨了应用标准组成技术来自计数平台的组成数据的负面影响,表明需要新的统计技术来解释生物样本中计数总数的差异。组合数据中关联的一个常见度量是“变异矩阵”,它包含组合中各部分成对对数比的方差。在基于计数的成分数据中,关于变异矩阵的估计缺乏统计理论。与此相关,零计数在微生物组数据中通常非常普遍。已经表明,关联和多样性的措施是敏感的,用于处理零计数的程序的选择。然而,贝叶斯非参数模型,如分层Pitman-Yor (HPY)过程,最近在模拟微生物组数据中的物种丰度分布方面被证明是成功的。我还将在HPY过程的背景下开发关联度量,这将允许在数据中存在零计数时估计关联。最后,我将研究这些方法在其他基于计数的组成平台(如单细胞RNA测序)上的适用性。这项工作将为估计宏基因组数据的关联和多样性提供新的方法,这些方法将解释生物样本中计数总数的差异以及零计数的丰度。虽然这个问题最近引起了人们的注意,但在这方面仍然缺乏可靠的统计理论。此外,这些方法将适用于许多其他类型的平台,如单细胞RNA测序数据集。
英文摘要
The human microbiome is the collection of organisms residing in or on the human body and plays an integral part in human health. A typical microbiome dataset consists of a count matrix where rows correspond to samples from different microbial populations and columns correspond to microbial taxa (e.g. species, genus, family). Microbiome sequencing data are compositional in nature, meaning that the count of a taxon in a sample should not be interpreted in absolute terms; rather, it should only be interpreted relative to one or more of the other taxa. This makes measuring associations between different taxa difficult, since correlations on the raw taxon proportions will be negatively biased. In this proposal I will extend existing multinomial model to estimate covariance matrices in compositional data as a function of covariates.  To do this, I will modify previous models I have developed to handle a more flexible covariance structure.  Relatedly, statistical techniques have been developed to address the pitfalls of compositional data, for example, the log-ratio transformation. However, the log-ratio transformation is scale invariant, whereas count-based sequencing platforms are not. Recently, (Lovell, Chua, and McGrath 2020) published an important paper exploring the negative effects of applying standard compositional techniques compositional data from count platforms, demonstrating a need for new statistical techniques that will account for differences in count totals across biological samples. A common measure of association in compositional data is the "variation matrix", which contains the variances of the pairwise log-ratios of the parts of the composition. There is a lack of statistical theory surrounding estimation of the variation matrix in count-based compositional data.  Relatedly, zero counts are often very prevalent in microbiome data. It has been shown that measures of association and diversity are sensitive to the choice of procedure used to handle zero counts. However, Bayesian non-parametric models such as the hierarchical Pitman-Yor (HPY) process have recently proved successful in modelling species abundance distributions in microbiome data.  I will additionally develop measures of association in the context of the HPY process, which will allow estimation of association when there are zero counts present in the data.  Finally, I will study the applicability of these methods on other count-based compositional platforms such as single-cell RNA sequencing. This work will provide new methods of estimating measures of association and diversity in metagenomic data that will account for differences in the count totals over biological samples as well as the abundance of zero counts. Though this problem has recently gained attention, there remains a lack of sound statistical theory in this context. Furthermore, these methods will be applicable to many other kinds of platforms such as single-cell RNA sequencing datasets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Measures of association in microbiome and other count-based platforms
  • 批准号:
    RGPIN-2021-03634
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2021
  • 负责人:
    McGregor, Kevin
  • 依托单位:
Measures of association in microbiome and other count-based platforms
  • 批准号:
    DGECR-2021-00459
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2021
  • 负责人:
    McGregor, Kevin
  • 依托单位:
国内基金
海外基金
精神分裂症全基因组关联研究的通路分析及验证
  • 批准号:
    81071087
  • 项目类别:
    面上项目
  • 资助金额:
    35.0万元
  • 批准年份:
    2010
  • 负责人:
    岳伟华
  • 依托单位:
精神分裂症与吸烟关联的分子遗传学机制研究
  • 批准号:
    81000579
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2010
  • 负责人:
    王志仁
  • 依托单位:
孤独症全基因组关联第二阶段研究
  • 批准号:
    81071110
  • 项目类别:
    面上项目
  • 资助金额:
    32.0万元
  • 批准年份:
    2010
  • 负责人:
    王力芳
  • 依托单位:
多盘科单殖吸虫宿主特异性及其与无尾两栖类宿主协同进化关系研究
  • 批准号:
    30960049
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2009
  • 负责人:
    范丽仙
  • 依托单位: