课题基金 / 基金详情

Improved Analysis of Metagenomes through the application of Read-Sized Profile HMMs to Marker Gene Subsequences

Improved Analysis of Metagenomes through the application of Read-Sized Profile HMMs to Marker Gene Subsequences
通过将 Read-Size Profile HMM 应用到标记基因子序列来改进宏基因组分析
批准号:
9181272
负责人:
Jeremy Selengut
金额:
$22.34万
依托单位国家:
美国
项目类别:
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-06-02 至 2018-05-31

项目摘要

项目成果

Jeremy Selengut的其他基金

相似基金

相关文献

中文摘要
翻译
项目摘要 对人类微生物群及其大量与宿主相关的生物体的研究表明, 极有希望增加我们对人类健康和疾病的了解。使用它的 与起源基因组信息无关的片段序列数据, 元基因组学面临的特殊挑战是如何提供可靠的功能 注释和分类赋值。 在这里,我们通过利用现有的轮廓隐马尔可夫模型来解决这些问题 (HMM)的功能特征基因家族。而不是依赖片段匹配来 全长基因或基因模型,我们将确定基因模型的哪些片段 能够对功能和起源进行高质量的注释,并专注于这些方面。通过 这种方法,基因模型中具有低序列保守性或具有 可变插入/间隙长度(倾向于低召回率),或由 多个基因家族和功能共享的序列(倾向于低精度) 被系统地消除,增加了总体信噪比。的高质量细分市场 模型(“迷你”HMM)将成为我们的分析工具。使用这些方法,我们希望提供 稳健的方法,将元基因组学从组装优先策略的限制中解放出来,以及 从而提供关于复合体中众多低丰度物种的信息 生物样本。 我们将使用细菌的单拷贝基因作为分类标记,并将产生一个 来自高质量基因组的这些基因的数据库。我们希望能找到大约80个合适的标记 基因,决定了几千个基因组。对于这些基因中的每一个,我们将产生一个 相应的参考系统发育树。在生产这些资源的过程中, 现有型号(TIGRFAM和Pfam HMM)将根据当前的 参考基因组和持续的、最先进的构建过程。这些资源,以及 我们生产的任何软件都将通过我们的公共网站提供。 有了这些方法和资源,我们将获得分类图谱,研究基因 并设计出将这些基因与剖面中的分类群联系起来的方法。我们将利用 真实的和合成的元基因组来执行方法的验证,并建立统计 我们结果的置信度指标。
英文摘要
Project Summary The study of the human microbiome, with its multitudes of host-associated organisms, holds great promise for increasing our understanding of human health and disease. With its fragmented sequence data unlinked from genome of origin information, the particular challenge of metagenomics is how to provide reliable functional annotation and taxonomic assignment. Here we address these issues by leveraging existing profile hidden Markov models (HMMs) of functionally characterized gene families. Instead of relying on fragment matches to full-length genes or gene models, we will determine which segments of gene models are capable of high-quality annotations of function and origin, and focus on those. By this approach, the portions of the gene models that have low sequence conservation or have variable insertion/gap length (tending towards low recall), or those that are composed of sequence shared among multiple gene families and functions (tending towards low precision) are systematically eliminated, increasing overall signal-to-noise. The high-quality segments of the models (“mini” HMMs) will be our analytical tools. Using these methods we hope to provide robust approach that frees metagenomics from the limitations of assembly-first strategies, and thereby provide access to information about the numerous low-abundance species in complex biological samples. We will use bacterial single-copy genes as taxonomic markers, and will produce a database of these genes from high-quality genomes. We expect to identify ~80 suitable marker genes, determined for several thousand genomes. For each of these genes, we will produce a corresponding reference phylogenetic tree. In the course of producing these resources, the existing models (TIGRFAMs and Pfam HMMs) will be updated based on the current set of reference genomes and a constant, state-of-the-art construction process. These resources, and any software we produce will be made available through our public website. With these methods and resources, we will obtain taxonomic profiles, investigate genes of interest and devise methods for linking those genes to the taxa in the profile. We will utilize real and synthetic metagenomes to perform validation of the methods, and establish statistical confidence metrics for our results.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Magnesium-dependent tyrosine phosphatases
海外基金