Improved Analysis of Metagenomes through the application of Read-Sized Profile HMMs to Marker Gene Subsequences
Improved Analysis of Metagenomes through the application of Read-Sized Profile HMMs to Marker Gene Subsequences
批准号:
9181272
负责人:
Jeremy Selengut
金额:
$22.34万
依托单位国家:
美国
项目类别:
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-06-02 至 2018-05-31
关键词:
Access to InformationAddressAmino Acid SequenceAntibiotic ResistanceBacteriaBacterial GenesBiologicalBiological ProcessComplexComputer softwareDataData AnalysesData SetDatabasesDetectionDiseaseDrug resistanceEventGene FamilyGene TransferGenesGenomeGenomic SegmentHealthHorizontal Gene TransferHumanHuman MicrobiomeInflammatoryKnowledgeLengthLinkMeasuresMetagenomicsMethodsModelingNatureNoiseOrganismOxidative StressPeptide Sequence DeterminationPerformancePhylogenetic AnalysisProceduresProcessReadingResourcesSamplingSignal TransductionTaxonTestingToxinTreesUpdateValidationWorkanalytical toolbasebiological adaptation to stressgene functiongenome sequencingimprovedinterestmarkov modelmembermetagenomemicrobiomenovel strategiesreference genomeresistance genetoolweb site
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Project Summary
The study of the human microbiome, with its multitudes of host-associated organisms, holds
great promise for increasing our understanding of human health and disease. With its
fragmented sequence data unlinked from genome of origin information, the
particular challenge of metagenomics is how to provide reliable functional
annotation and taxonomic assignment.
Here we address these issues by leveraging existing profile hidden Markov models
(HMMs) of functionally characterized gene families. Instead of relying on fragment matches to
full-length genes or gene models, we will determine which segments of gene models are
capable of high-quality annotations of function and origin, and focus on those. By
this approach, the portions of the gene models that have low sequence conservation or have
variable insertion/gap length (tending towards low recall), or those that are composed of
sequence shared among multiple gene families and functions (tending towards low precision)
are systematically eliminated, increasing overall signal-to-noise. The high-quality segments of
the models (“mini” HMMs) will be our analytical tools. Using these methods we hope to provide
robust approach that frees metagenomics from the limitations of assembly-first strategies, and
thereby provide access to information about the numerous low-abundance species in complex
biological samples.
We will use bacterial single-copy genes as taxonomic markers, and will produce a
database of these genes from high-quality genomes. We expect to identify ~80 suitable marker
genes, determined for several thousand genomes. For each of these genes, we will produce a
corresponding reference phylogenetic tree. In the course of producing these resources, the
existing models (TIGRFAMs and Pfam HMMs) will be updated based on the current set of
reference genomes and a constant, state-of-the-art construction process. These resources, and
any software we produce will be made available through our public website.
With these methods and resources, we will obtain taxonomic profiles, investigate genes
of interest and devise methods for linking those genes to the taxa in the profile. We will utilize
real and synthetic metagenomes to perform validation of the methods, and establish statistical
confidence metrics for our results.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Magnesium-dependent tyrosine phosphatases
-
批准号:6109157
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Jeremy Selengut
-
依托单位:
海外基金