Estimating evolutionary and demographic parameters via ARG-derived IBD.

Estimating evolutionary and demographic parameters via ARG-derived IBD.
复制标题

通过 ARG 衍生的 IBD 估计进化和人口统计参数。

DOI:
10.1101/2024.03.07.583855
复制
发表时间:
2024
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Balding,DavidJ
Balding,DavidJ
中科院分区:
--
文献类型:
--
作者:
Huang,Zhendong;Kelleher,Jerome;Chan,Yao-Ban;Balding,DavidJ

文献摘要

相似文献

从基因组序列的样本推断进化和人口统计学参数通常通过首先推断血统相同(IBD)基因组区段来进行。通过利用基于祖先重组图(ARG)的有效数据编码,我们获得了优于当前方法的三个主要优点:(i)不需要对IBD片段施加长度阈值,(ii)可以在没有难以验证的无重组要求的情况下定义IBD,以及(iii)仅使用线性缩放的一组序列对中的IBD片段,可以在几乎不损失统计效率的情况下减少计算时间样本大小。我们首先展示了强大的推论时,真正的IBD信息是从模拟数据。对于从真实的数据推断的IBD,我们提出了一个近似贝叶斯计算推断算法,并使用它来表明,即使是推断不佳的短IBD段可以提高估计。尽管用于推断的数据减少了4000倍,但我们的突变率估计器达到了与先前发表的方法相似的精度,并且我们确定了人群之间的显著差异。在我们的方法中,计算成本限制了模型的复杂性,但我们能够将未知的滋扰参数和模型误指定,仍然找到改进的参数推断。
Inference of evolutionary and demographic parameters from a sample of genome sequences often proceeds by first inferring identical-by-descent (IBD) genome segments. By exploiting efficient data encoding based on the ancestral recombination graph (ARG), we obtain three major advantages over current approaches: (i) no need to impose a length threshold on IBD segments, (ii) IBD can be defined without the hard-to-verify requirement of no recombination, and (iii) computation time can be reduced with little loss of statistical efficiency using only the IBD segments from a set of sequence pairs that scales linearly with sample size. We first demonstrate powerful inferences when true IBD information is available from simulated data. For IBD inferred from real data, we propose an approximate Bayesian computation inference algorithm and use it to show that even poorly-inferred short IBD segments can improve estimation. Our mutation-rate estimator achieves precision similar to a previously-published method despite a 4 000-fold reduction in data used for inference, and we identify significant differences between human populations. Computational cost limits model complexity in our approach, but we are able to incorporate unknown nuisance parameters and model misspecification, still finding improved parameter inference.