课题基金 / 基金详情

Mathematical foundations of non-reversible MCMC for genome-scale inference

Mathematical foundations of non-reversible MCMC for genome-scale inference
用于基因组规模推理的不可逆 MCMC 的数学基础
批准号:
EP/V049208/1
负责人:
Jere Koskela
金额:
$9.71万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
已结题
起止时间:
2021 至 --

项目摘要

项目成果

Jere Koskela的其他基金

相似基金

相关文献

中文摘要
翻译
世界正在经历基因DNA序列数据的爆炸式增长。数据集中的模式包含了关于不可观察的种群生物和人口历史的信息,这些信息反过来又推动了医学、人口统计学和自然保护等领域的发现。将观察到的模式与可测试的预测和推断联系起来的一个中心工具是祖先重组图,它沿着采样的DNA序列对共同祖先的模式进行建模。由于共同祖先通常不能直接观察到,推论是通过对可能的祖先进行平均来进行的。在非常简单的情况下,可以精确地进行平均,但在生物学相关的情况下,通常必须近似。典型的近似方法创建候选祖先的集合,并使用集合平均值作为真实平均值的代理。这个过程的准确性取决于集合在多大程度上代表了所有可能祖先的集合。保证具有代表性的集成所需的计算时间随着数据集规模的增加而迅速增长,并且在实践中,这种基于集成的方法只能应用于现代标准的小型数据集。在实践中,研究人员采用计算速度更快的方法,其理论性能尚不清楚。缺乏理论基础会使研究结果的可解释性复杂化,并使其难以准确量化其相关的不确定性。在过去的几年里,一种新的用于构建具有代表性的集成的方法被称为锯齿形算法,已经被开发出来并变得越来越普遍。它在遗传学的试点应用中也显示出前景,但将之字形方法应用于基因组规模数据的有效数据结构是必不可少的组成部分,目前尚不清楚。该项目旨在开发和测试合适的数据结构,使软件包的工程能够将可行的运行时间与数据集中统计信号的有效使用结合起来。
英文摘要
The world is undergoing an explosion of genetic DNA sequence data. Patterns within data sets carry information about unobservable biological and demographic histories of populations, which in turn are fueling discoveries in areas such as medicine, demography, and conservation. A central tool connecting observed patterns to testable predictions and inference is the Ancestral Recombination Graph, which models the patterns of common ancestry along sampled DNA sequences. Since common ancestry is typically not observable directly, inferences are made by averaging over possible ancestries. In very simple cases the averaging can be carried out exactly, but in biologically relevant settings it typically has to be approximated. Typical approximation methods create an ensemble of candidate ancestries, and use the ensemble average as a proxy for the true average. The accuracy of this procedure depends on the degree to which the ensemble is representative of the set of all possible ancestries. The computational time required to guarantee a representative ensemble grows rapidly as the size of a data set increases, and in practice, such ensemble-based methods can only be applied to small data sets by modern standards. In practice, researchers resort to computationally faster methods, the theoretical performance of which is less well understood. The lack of theoretical foundations can complicate the interpretability of findings, and makes it difficult to accurately quantify their associated uncertainty.A new class of methods for building representative ensembles, called zig-zag algorithms, has been developed and become increasingly widespread over the last several years. It has also shown promise in pilot applications in genetics, but an effective data structure for applying zig-zag methods to genome-scale data is an essential ingredient, and remains unknown. This project aims to develop and test suitable data structures, making possible the engineering of software packages which combine feasible run times with efficient use of statistical signal in data sets.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1093/bioinformatics/btad017
发表时间: 2023-01-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: []
通讯作者:
Weak convergence of non-neutral genealogies to Kingman's coalescent
非中立谱系与金曼合并的弱收敛
DOI: 10.1016/j.spa.2023.04.016
发表时间: 2023
期刊: Stochastic Processes and their Applications
影响因子: 1.4
作者: [Brown S]
通讯作者: Brown S
DOI: 10.7554/elife.80781
发表时间: 2023-02-20
期刊: eLife
影响因子: 7.7
作者: [Árnason E, Koskela J, Halldórsdóttir K, Eldon B]
通讯作者: Eldon B
DOI: 10.1101/2022.05.29.493887
发表时间: 2022-12
期刊: eLife
影响因子: 7.7
作者: [E. Árnason;Jere Koskela;Katrín Halldórsdóttir;Bjarki Eldon]
通讯作者: E. Árnason;Jere Koskela;Katrín Halldórsdóttir;Bjarki Eldon
共 7 条
    Exact scalable inference for coalescent processes
    • 批准号:
      EP/R044732/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $12.68万
    • 财政年份:
      2018
    • 负责人:
      Jere Koskela
    • 依托单位:
    海外基金