课题基金 / 基金详情

ABI Innovation: Posterior Predictive Checks of Evolutionary Models.

ABI Innovation: Posterior Predictive Checks of Evolutionary Models.
ABI 创新:进化模型的后验预测检查。
批准号:
1661029
负责人:
Bryan Carstens
金额:
$40.36万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-05-15 至 2021-09-30

项目摘要

项目成果

Bryan Carstens的其他基金

相似基金

相关文献

中文摘要
翻译
在生物基础设施部生物信息学进展的支持下,俄亥俄州立大学的Bryan Carstens教授和他的研究小组将进一步开发P2M2软件包,以便它可以更好地估计多重事件如何影响群体遗传学。生物之间的大多数差异最终都可以追溯到基因变异,包括人类。由于种群规模的变化、种群间的迁移或长时间的隔离等事件,一个物种的遗传差异模式在不同的种群中出现。例如,了解如何在具有迁移限制的一定规模的自然种群中建立有益的突变,对于了解人类健康、保护遗传学、农业育种和许多其他研究领域的各个方面都很重要。最近的技术进步使得从许多人身上获得基因序列变得快速而廉价,假设样本可以收集。许多软件包提供了模型,可以根据当前的遗传数据(包括种群规模或迁移率等生物参数)来估计先前事件的样子。然而,每一种方法都假定了一个特定的人口统计学数学模型,并且仅限于估计可能影响遗传变异的生物过程子集的参数。分析模型的假设与真实种群历史之间的不匹配将产生不准确的参数估计,这可能会误导生物学推断。该项目将开发软件,使生物学家能够评估特定软件包对给定遗传数据集的适用性。因此,它将通过提高从遗传数据中得出的生物推断的质量,从保护濒危物种的努力到调查病毒病原体的历史,从而造福社会。贝叶斯推理通常用于分析遗传数据,因为它提供了一种计算效率高的方法来识别参数空间的高概率区域,但所有这些推断都取决于选择在分析中使用的模型。虽然现有的分析模型可以估计与所有种群水平的生物过程相关的参数,如遗传漂变、系统发育分化、基因流动、种群大小变化等,但计算限制使任何给定的分析模型都无法纳入少数这些过程。生物学家通常会凭直觉选择使用哪种分析方法,并且通常缺乏评估给定遗传数据的模型的绝对统计拟合的方法。因此,从遗传数据分析得出的推论实际上取决于用于分析数据的模型的适当性,尽管它们很少以这种方式提出。该研究将对P2C2M R包进行扩展,该包目前实现了后验预测模拟,以评估单一模型(多物种聚结模型)的统计拟合。这项工作将扩展P2C2M,以便可以评估其他聚结方法的统计拟合。通过扩展P2C2M,本研究将模型拟合作为从遗传数据进行生物学推断的整个过程中的一个重要步骤。生物学家已经投入了大量的精力来证明他们用来用口头推理和定性论证来分析数据的模型是正确的,但通常缺乏以直接定量方式进行分析的工具和统计框架。P2C2M将在项目完成时提供这些工具。扩展后的P2C2M R包的直接结果是,进化遗传学家做出的推论将更有见地,因为研究人员和他们的读者在选择分析模型时将更有信心,这些模型是推导这些推论的依据。这项工作将增强具有重要社会效益的生物学推断,例如鉴定隐物种、了解入侵物种和疾病媒介的人口统计学,以及濒危物种的等位基因在景观中的移动。更新的项目代码将在https://cran.r-project.org/web/packages/P2C2M/index.html上提供,其他补充信息将在https://carstenslab.osu.edu/上发布。
英文摘要
With support from the Advances in Biological Informatics in the Division of Biological Infrastructure, Professor Bryan Carstens and his research group at the Ohio State University will further develop the P2M2 software package, so that it can better estimate how multiple events may shape population genetics. Most of the differences among living organisms ultimately can be traced to genetic variation, including humans. Patterns of genetic differences in a species appear across different groups, due to events such as changes in population size, migration between populations, or long periods of isolation. Understanding, for example, how beneficial mutations can be established in a natural population of a certain size that has migration limits is important for understanding aspects of human health, conservation genetics, breeding in agriculture and many other research areas. Recent technological advances have made it fast and inexpensive to obtain genetic sequences from many individuals, assuming samples can be collected. A number of software packages provide models that estimate what prior events looked like from current genetic data, including biological parameters such as population size or migration rate. However, each assumes a particular mathematic model of population demography and are limited to estimating parameters from a subset of the biological processes that may influence genetic variation. A poor match between the assumptions of the analytical model and the true population history will produce inaccurate parameter estimates that are likely to mislead the biological inference. This project will develop software that enables biologists to assess how appropriate a particular software package is to a given set of genetic data. Therefore, it will benefit society by improving the quality of biological inferences drawn from genetic data, ranging from efforts to protect endangered species to investigations into the history of viral pathogens.Bayesian inference is commonly used to analyze genetic data because it provides a computationally efficient approach to identifying highly-probable regions of parameter space, but all such inference is conditional on the models chosen to use in the analysis. While analytical models exist that can estimate parameters associated with all population-level biological processes, such as genetic drift, phylogenetic divergence, gene flow, population size change, etc., computational limitations prevent any given analytical model of incorporating more than a handful of these processes. Biologists typically choose which analytical method to use intuitively, and generally lack approaches for assessing the absolute statistical fit of a model given the genetic data. Consequently, the inferences that result from the analysis of genetic data are effectively conditional on the appropriateness of the model used to analyze the data, although they are rarely presented in such terms. The proposed work will develop and implement a considerable expansion of the P2C2M R package, which currently implements posterior predictive simulation to assess the statistical fit of a single model - the multispecies coalescent model. The work will expand P2C2M such that the statistical fit of additional coalescent methods can be evaluated. By expanding P2C2M, the work promotes the consideration of model fit as an important step within the overall process of making biological inferences from genetic data. Biologists have devoted a great deal of energy to justify the models that they use to analyze their data using verbal reasoning and qualitative arguments, but have generally lacked the tools and statistical framework to do so in a direct quantitative manner. P2C2M will provide these tools by the time of project completion. As a direct consequence of the expanded P2C2M R package, the inferences made by evolutionary geneticists will be more insightful and because researchers and their audiences will have enhanced confidence in the choice of analytical models from which these inferences are derived. The work will enhance biological inferences with important societal benefit, such as the identification of cryptic species, understanding the demography of invasive species and disease vectors, and the movement of alleles across the landscape in endangered species. Updated project code will be available at https://cran.r-project.org/web/packages/P2C2M/index.html and other supplemental information distributed at https://carstenslab.osu.edu/.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.7717/peerj.8271
发表时间: 2020-01-10
期刊: PEERJ
影响因子: 2.7
作者: [Duckett, Drew J., Pelletier, Tara A., Carstens, Bryan C.]
通讯作者: Carstens, Bryan C.
ICBR Capacity: Biological Collections: Infrastructure improvement and data preservation of the Tetrapods Collection at the Ohio State University Museum of Biological Diversity.
  • 批准号:
    2312986
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $47.45万
  • 财政年份:
    2023
  • 负责人:
    Bryan Carstens
  • 依托单位:
SG: Leveraging massive song databases and deep learning to examine the mechanisms causing diversification of bird vocalizations.
  • 批准号:
    2016189
  • 项目类别:
    Standard Grant
  • 资助金额:
    $19.98万
  • 财政年份:
    2020
  • 负责人:
    Bryan Carstens
  • 依托单位:
Collaborative Research:Aggregating and Repurposing Phylogeographic Data.
  • 批准号:
    1910623
  • 项目类别:
    Standard Grant
  • 资助金额:
    $50.33万
  • 财政年份:
    2019
  • 负责人:
    Bryan Carstens
  • 依托单位:
Dimensions US-BIOTA-Sao Paulo: Traits as predictors of adaptive diversification along the Brazilian Dry Diagonal.
  • 批准号:
    1831319
  • 项目类别:
    Standard Grant
  • 资助金额:
    $38.31万
  • 财政年份:
    2018
  • 负责人:
    Bryan Carstens
  • 依托单位:
海外基金