Sibship reconstruction from genetic data with typing errors

Sibship reconstruction from genetic data with typing errors
复制标题

DOI:
10.1534/genetics.166.4.1963
复制
发表时间:
2004-04-01
期刊:
影响因子:
3.3
通讯作者:
Wang, JL
Wang, JL
中科院分区:
生物学2区
文献类型:
--
作者:
Wang, JL

文献摘要

被引文献

相似文献

已经开发了似然法,利用没有父母信息的遗传标记数据,将样本中的个体划分为全同胞和半同胞家庭。他们总是做出关键的假设,即标记数据没有基因分型错误和突变,因此在推断兄弟姐妹关系方面是完全可靠的。然而,不幸的是,这一假设很少在实际应用中适用于几乎所有类型的遗传标记,如果违反,可能会严重偏离亲缘关系估计,如本文的模拟所示。我提出了一种新的似然方法,其中包含了简单而稳健的打字错误模型。仿真结果表明,该方法可以从错误率较高的标记数据中准确地推断出全同胞关系和半同胞关系,并能在每个重构的同胞家系中识别每个基因座的打字错误。新方法还改进了以往的方法,采用了一种新的迭代过程来更新等位基因频率,并考虑了重建的亲缘关系,允许使用亲本信息,并使用了计算似然函数和搜索最大似然配置的有效算法。对不同数量的标记基因座、不同的错误率、不同的样本大小和家族结构进行了广泛的石油模拟数据测试,并应用于两个经验数据集以验证其有效性。
Likelihood methods have been developed to partition individuals in a sample into full-sib and half-sib families using genetic marker data without parental information. They invariably make the critical assumption that marker data are free of genotyping errors and mutations and are thus completely reliable in inferring sibships. Unfortunately, however, this assumption is rarely tenable for virtually all kinds of genetic markers in practical use and, if violated, can severely bias sibship estimates as shown by simulations in this article. I propose a new likelihood method With simple and robust models of typing error incorporated into it. Simulations show that the new method call be used to infer full- and half-sibships accurately from marker data with a high error rate and to identify typing errors at each locus in each reconstructed sib family. The new method also improves previous ones by adopting a Fresh iterative procedure for updating allele frequencies with reconstructed sibships taken into account, by allowing for the use of parental information, and by using efficient algorithms for calculating the likelihood function and searching for the maximum-likelihood configuration. It is tested extensively oil simulated data with a varying number of marker loci, different rates of typing errors, and various sample sizes and family structures and applied to two empirical data sets to demonstrate its usefulness.