The power of single-nucleotide polymorphisms for large-scale parentage inference

The power of single-nucleotide polymorphisms for large-scale parentage inference
复制标题

DOI:
10.1534/genetics.105.048074
复制
发表时间:
2006-04-01
期刊:
影响因子:
3.3
通讯作者:
Garza, JC
Garza, JC
中科院分区:
生物学2区
文献类型:
--
作者:
Anderson, EC;Garza, JC

文献摘要

被引文献

相似文献

基于似然的亲子关系推断取决于似然比统计量的分布,在大多数感兴趣的情况下,不能准确确定,而只能通过蒙特卡罗模拟近似。我们提供的重要性抽样算法,有效地近似非常小的尾部概率分布的似然比统计。这些重要性抽样方法允许估计所有假阳性率,因此允许在涉及大量潜在父母和许多潜在后代的大型研究中基于似然性的亲子关系推断。我们研究了这些重要性抽样算法在使用单核苷酸多态性(SNP)数据进行亲子关系推断的情况下的性能,发现它们可以加速尾部概率的计算> 100万倍。随后,我们使用重要性抽样算法来计算大规模亲子关系研究中SNP的功效,特别注意基因分型错误的影响和假定的母亲-父亲-后代三人组成员中相关个体的发生。这些模拟表明,60-100个SNP可能允许准确的谱系重建,即使在涉及数千个潜在母亲、父亲和后代的情况下。此外,我们比较的权力排除为基础的亲子关系推断的可能性为基础的方法。在许多情况下,基于可能性的推断更加强大;基于排除的推断将需要多40%的SNP位点,以实现与基于可能性的方法相同的准确性。我们的研究结果表明,单核苷酸多态性是一个强大的工具,在大型管理和/或自然人群的亲子关系推断。
Likelihood-based parentage inference depends on the distribution of a likelihood-ratio statistic, which, in most cases of interest, cannot be exactly determined, but only approximated by Monte Carlo simulation. We provide importance-sampling algorithms for efficiently approximating very small tail probabilities in the distribution of the likelihood-ratio statistic. These importance-sampling methods allow the estimation of sin all false-positive rates and hence permit likelihood-based inference of parentage in large studies involving a great number of potential parents and many potential offspring. We investigate the performance of these importance-sampling algorithms in the context of parentage inference using single-nucleotide polymorphism (SNP) data and find that they may accelerate the computation of tail probabilities > 1 millionfold. We subsequently use the importance-sampling algorithms to calculate the power available with SNPs for largescale parentage studies, paying particular attention to the effect of genotyping errors and the occurrence of related individuals among the members of the putative mother-father-offspring trios. These simulations show that 60-100 SNPs may allow accurate pedigree reconstruction, even in situations involving thousands of potential mothers, fathers, and offspring. In addition, we compare the power of exclusion-based parentage inference to that of the likelihood-based method. Likelihood-based inference is much more powerful under many conditions; exclusion-based inference would require 40% more SNP loci to achieve the same accuracy as the likelihood-based approach in one common scenario. Our results demonstrate that SNPs are a powerful tool for parentage inference in large managed and/or natural populations.