Bayesian Joint Estimation of Alignment and Phylogeny
Bayesian Joint Estimation of Alignment and Phylogeny
批准号:
7660485
负责人:
Marc A. Suchard
金额:
$29.87万
依托单位国家:
美国
项目类别:
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-08-01 至 2013-07-31
关键词:
AddressAlgorithmsBioinformaticsBiologyBiomedical ResearchCommunicable DiseasesComparative StudyDataData SetEducational process of instructingEvolutionFosteringGenesGeneticGenomicsHeterogeneityHumanJointsKnowledgeLifeMathematicsMethodsModelingMolecularMolecular BiologyMotivationMutateOrganismPhylogenetic AnalysisPhylogenyPlayProcessPropertyProtein FamilyRoleSequence AlignmentSequence AnalysisSequence HomologsSideSisterSpecific qualifier valueTaxonTechniquesTestingTimeTrainingTreesVariantViruscomparativeconditioningimprovedinsertion/deletion mutationinterestlife historymarkov modelnovelreconstructiontooluser friendly software
中文摘要
描述(由申请人提供):系统发育重建是研究分子序列的宝贵工具。从描述序列中的字符如何随时间变化开始,这些方法试图揭示序列的相关性。常见的应用范围从进化生物学中描述生物体的进化历史到分子生物学和生物信息学中估计遗传距离和构建蛋白质家族。标准的重建方法依赖于序列比对,该序列比对指定序列中的哪些特征是同源的,源自共同的祖先。一个根本的困难是序列比对不能直接观察到,它们是原始序列数据的推断性质,必须沿着估计序列发生。目前的工具顺序地处理这种推断,首先确定对比对的有时较差的估计,然后根据比对的真实性来重建同源性。该项目为最终用户提供了实用的工具,可以同时推断序列估计引入的对齐和重复性,侧步偏差。这些工具假设字符替换模型和插入/删除(indel)过程,通过插入/删除过程添加或删除字符以生成对齐。此外,这些插入缺失从数据中提供先前未充分利用的信息以推断植物发生。主要的进步使这个对齐框架对现实生活中的数据集很有用。该框架在很大程度上借鉴了隐马尔可夫模型,贝叶斯计算和巧妙的参数集成,以产生一个计算效率高的推理引擎。专家的先验知识有助于为indel过程提供信息。由此,现实的先验使贝叶斯因子检验能够解决特定的插入缺失是由血统共享还是同源的,从而减少了关于它们在遗传学中的价值的争议。建模假设更好地反映了潜在的生物学。在插入缺失过程中允许空间变化提供了更准确的植物发生和比对。扩展还提供了异质性测试,以确定进化感兴趣的序列区域。这些方法的例子跨越了进化的所有时间尺度,从数十亿年来推断生命之树的早期分支到几个月来描述受感染宿主内快速进化的病毒的多样化。
该项目对生物医学研究的许多领域产生了显著影响。例如,该项目在生物信息学中提供数学和统计培训,这将在21世纪的发现中发挥主要作用,采用双序列比对的严格推理工具提供改进的分子比较研究,更准确地了解人类进化和对抗传染病的新视角。
英文摘要
DESCRIPTION (provided by applicant): Phylogenetic reconstruction is an invaluable tool for studying molecular sequences. Starting from a description of how the characters in the sequences mutate over time, the methods attempt to uncover the sequences' relatedness. Common applications range from describing the evolutionary histories of living organisms in evolutionary biology to estimating genetic distances and constructing protein families in molecular biology and bioinformatics. Standard reconstruction methods rely on sequence alignments that specify which characters in the sequences are homologous, deriving from common ancestors. A fundamental difficulty is that sequence alignments are not directly observed; they are inferred properties of the raw sequence data and must be estimated along with the phylogeny. Current tools handle this inference sequentially, first determining a sometimes poor estimate of the alignment and then conditioning on the truth of alignment to reconstruct the phylogeny. This project provides practical tools for end-users to simultaneously infer alignment and phylogeny, side-stepping biases that sequential estimation introduces. The tools assume both a character substitution model and an insertion/deletion (indel) process through which characters are added or removed generating an alignment. Further, these indels supply previously under-utilized information from the data to infer phytogenies. Major advances make this phylo-alignment framework useful for real-life datasets. The framework draws heavily on hidden Markov models, Bayesian computation and clever parameter integration to produce a computationally efficient inference engine. Expert prior knowledge helps inform the indel process. From this, realistic priors enable Bayes factor tests to address if specific indels are shared by descent or are homoplastic, reducing controversy over their value in phylogenetics. Modeling assumptions better reflect the underlying biology. Allowing spatial variation in the indel process provides more accurate phytogenies and alignments. The extensions also provide for heterogeneity tests to identify evolutionary interesting sequence regions. Examples of the methods span all time-scales of evolution, across billions of years to infer early branches in the Tree of Life to matters of months to describe the diversification of rapidly evolving viruses within infected hosts.
This project markedly impacts many fields across biomedical research. For example, the project furnishes mathematical and statistical training in bioinformatics which will play a prime role in discovery during the 21st century, and rigorous inference tools employing phylo-alignment deliver improved molecular, comparative studies, a more accurate understanding of human evolution and new perspectives from which to battle infectious diseases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical innovation to integrate sequences and phenotypes for scalable phylodynamic inference
-
批准号:10584588
-
项目类别:
-
资助金额:$45.9万
-
财政年份:2021
-
负责人:Marc A. Suchard
-
依托单位:
Statistical innovation to integrate sequences and phenotypes for scalable phylodynamic inference
-
批准号:10390334
-
项目类别:
-
资助金额:$46.59万
-
财政年份:2021
-
负责人:Marc A. Suchard
-
依托单位:
Statistical innovation to integrate sequences and phenotypes for scalable phylodynamic inference
-
批准号:10177121
-
项目类别:
-
资助金额:$47.83万
-
财政年份:2021
-
负责人:Marc A. Suchard
-
依托单位:
Consortium for Viral Systems Biology Modeling Core
-
批准号:10579085
-
项目类别:
-
资助金额:$7.5万
-
财政年份:2018
-
负责人:Marc A. Suchard
-
依托单位:
Consortium for Viral Systems Biology Modeling Core
-
批准号:10374718
-
项目类别:
-
资助金额:$42.5万
-
财政年份:2018
-
负责人:Marc A. Suchard
-
依托单位:
Consortium for Viral Systems Biology Modeling Core
-
批准号:10310604
-
项目类别:
-
资助金额:$5.77万
-
财政年份:2018
-
负责人:Marc A. Suchard
-
依托单位:
Bayesian Joint Estimation of Alignment and Phylogeny
-
批准号:7596504
-
项目类别:
-
资助金额:$30.32万
-
财政年份:2008
-
负责人:Marc A. Suchard
-
依托单位:
Bayesian Joint Estimation of Alignment and Phylogeny
-
批准号:8116012
-
项目类别:
-
资助金额:$29.54万
-
财政年份:2008
-
负责人:Marc A. Suchard
-
依托单位:
Bayesian Joint Estimation of Alignment and Phylogeny
-
批准号:7883433
-
项目类别:
-
资助金额:$29.83万
-
财政年份:2008
-
负责人:Marc A. Suchard
-
依托单位:
Bayesian Joint Estimation of Alignment and Phylogeny
-
批准号:8302280
-
项目类别:
-
资助金额:$29.52万
-
财政年份:2008
-
负责人:Marc A. Suchard
-
依托单位:
海外基金