Mathematical foundations of non-reversible MCMC for genome-scale inference
Mathematical foundations of non-reversible MCMC for genome-scale inference
批准号:
EP/V049208/1
负责人:
Jere Koskela
金额:
$9.71万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
已结题
起止时间:
2021 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The world is undergoing an explosion of genetic DNA sequence data. Patterns within data sets carry information about unobservable biological and demographic histories of populations, which in turn are fueling discoveries in areas such as medicine, demography, and conservation. A central tool connecting observed patterns to testable predictions and inference is the Ancestral Recombination Graph, which models the patterns of common ancestry along sampled DNA sequences. Since common ancestry is typically not observable directly, inferences are made by averaging over possible ancestries. In very simple cases the averaging can be carried out exactly, but in biologically relevant settings it typically has to be approximated. Typical approximation methods create an ensemble of candidate ancestries, and use the ensemble average as a proxy for the true average. The accuracy of this procedure depends on the degree to which the ensemble is representative of the set of all possible ancestries. The computational time required to guarantee a representative ensemble grows rapidly as the size of a data set increases, and in practice, such ensemble-based methods can only be applied to small data sets by modern standards. In practice, researchers resort to computationally faster methods, the theoretical performance of which is less well understood. The lack of theoretical foundations can complicate the interpretability of findings, and makes it difficult to accurately quantify their associated uncertainty.A new class of methods for building representative ensembles, called zig-zag algorithms, has been developed and become increasingly widespread over the last several years. It has also shown promise in pilot applications in genetics, but an effective data structure for applying zig-zag methods to genome-scale data is an essential ingredient, and remains unknown. This project aims to develop and test suitable data structures, making possible the engineering of software packages which combine feasible run times with efficient use of statistical signal in data sets.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1093/bioinformatics/btad017
发表时间:
2023-01-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
[]
通讯作者:
DOI:
10.1016/j.spa.2023.04.016
发表时间:
2023
期刊:
Stochastic Processes and their Applications
影响因子:
1.4
作者:
[Brown S]
通讯作者:
Brown S
DOI:
10.7554/elife.80781
发表时间:
2023-02-20
期刊:
eLife
影响因子:
7.7
作者:
[Árnason E, Koskela J, Halldórsdóttir K, Eldon B]
通讯作者:
Eldon B
DOI:
10.1101/2022.05.29.493887
发表时间:
2022-12
期刊:
eLife
影响因子:
7.7
作者:
[E. Árnason;Jere Koskela;Katrín Halldórsdóttir;Bjarki Eldon]
通讯作者:
E. Árnason;Jere Koskela;Katrín Halldórsdóttir;Bjarki Eldon
Bernoulli factories and duality in Wright-Fisher and Allen-Cahn models of population genetics
伯努利工厂以及群体遗传学赖特-费舍尔和艾伦-卡恩模型中的二元性
DOI:
10.1016/j.tpb.2024.01.002
发表时间:
2024
期刊:
Theoretical Population Biology
影响因子:
1.4
作者:
[Koskela J]
通讯作者:
Koskela J
共 7 条
Exact scalable inference for coalescent processes
-
批准号:EP/R044732/1
-
项目类别:Research Grant
-
资助金额:$12.68万
-
财政年份:2018
-
负责人:Jere Koskela
-
依托单位:
海外基金