Mathematical and computational analysis for species tree inference
Mathematical and computational analysis for species tree inference
批准号:
9295036
负责人:
John Rhodes
金额:
$38.7万
依托单位国家:
美国
项目类别:
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-08-01 至 2019-04-30
关键词:
AddressAlgebraAnimal ModelBayesian AnalysisBehaviorBiologyCombinatoricsComputer AnalysisComputer softwareDataData SetFoundationsGenesGeneticGenomicsGoalsKnowledgeMathematicsMethodologyMethodsModelingOrganismPhylogenetic AnalysisProbabilityPropertySamplingScienceStatistical ModelsTechniquesTreesVariantWorkbasecombinatorialexperimental studygenomic dataimprovedinsightmathematical analysispathogenpublic health relevancesimulationstatisticstool
中文摘要
描述(由申请人提供):了解生物体之间的进化关系是生物学中各种问题的基础。该项目研究和开发从遗传数据推断物种关系的新方法,利用以物种树为条件的基因树概率模型。它的主要目标是(1)推进这些模型的数学理解,着眼于物种树推断;(2)通过考虑来自基因树的新的和未充分利用的数据类型,包括分支、分裂、无根基因树和排序基因树,开发改进的物种树推断方法;(3)验证这些新方法的理论、计算和统计特性;(4)制作软件供经验生物学家使用。该项目将确定基因树汇总统计数据,并将利用这些统计数据开发实用方法,可用于缺失数据和违反模型假设的情况。将研究新方法和当前方法的数学,统计和计算特性,以进行比较,从而指导经验应用。基于模型的,概率的方法,这项工作提供了一个基础,以提高物种树推断基因树样本,从而从遗传序列数据。该项目解决了计算密集型全似然和贝叶斯分析之间的一个有前途的方法中间地带,这对于基因组规模的数据集来说往往是不可行的,而易于处理的组合方法往往缺乏理想的统计行为。这项工作将通过分析汇总统计量的行为,加深对基因树不一致的概率模型的了解,从而推进系统发育分析。它将通过引入新的统计一致的方法和通过发展对方法的鲁棒性的理论和实验理解来改进物种树推断的实践。此外,它使用的数学技术,从概率,组合学,代数统计,以及计算实验采用模拟,将提高数学进化生物学更普遍。
英文摘要
DESCRIPTION (provided by applicant): Understanding the evolutionary relationships between organisms is fundamental in a wide variety of problems in biology. This project investigates and develops new methods for inferring species relationships from genetic data, utilizing probabilistic models of gene trees conditional on a species tree. Its main goals are (1) to advance the mathematical understanding of these models, with a view toward species tree inference; (2) to develop improved methods for species tree inference by considering new and underutilized data types derived from gene trees, including clades, splits, unrooted gene trees, and ranked gene trees; (3) to validate theoretical, computational, and statistical properties of these new methods; (4) to produce software for use by empirical biologists. The project will identify gene tree summary statistics on which accurate inference can be based, and will employ these statistics to develop practical methods that can be used in the presence of missing data and under violations of model assumptions. The mathematical, statistical, and computational properties of both new and current methods will be studied to enable comparisons that can guide empirical applications. The model-based, probabilistic approach of this work provides a foundation for enhancing species tree inference from gene tree samples, and thus from genetic sequence data. The project addresses a promising methodological middle ground between computationally intensive full likelihood and Bayesian analyses, which are often infeasible for genomic-scale data sets, and tractable combinatorial methods, which often lack desirable statistical behaviors. The work will advance phylogenetic analysis by deepening knowledge of probabilistic models of gene tree discordance through analysis of the behavior of summary statistics. It will improve the practice of species tree inference by introducing new statistically consistent approaches and by developing theoretical and experimental understanding of the robustness of methods. Further, its use of mathematical techniques from probability, combinatorics, and algebraic statistics, as well as computational experiments employing simulation, will enhance mathematical evolutionary biology more generally.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金