Accurate partition of individuals into full-sib families from genetic data without parental information.

Accurate partition of individuals into full-sib families from genetic data without parental information.
复制标题

在没有父母信息的情况下,根据遗传数据将个体准确划分为全同胞家庭。

DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
3.3
通讯作者:
Heather R. Merry
Heather R. Merry
中科院分区:
生物学2区
文献类型:
--
作者:
Bruce Smith;C. Herbinger;Heather R. Merry

文献摘要

参考文献

被引文献

相似文献

提出了两种马尔可夫链蒙特卡罗算法,允许在没有亲代信息的情况下使用单位点遗传标记数据将个体划分为全同胞群体。这些算法提出了一种通过兄弟姐妹配置空间移动的方法,并在全兄弟姐妹或不相关的两两似然比的基础上找到最大化总分的配置,或者最大化所提议的家庭结构的全部联合似然。使用这些方法,759条大西洋鲑鱼中有757条使用4个微卫星标记被正确地划分为12个大小不等的全同胞家族。我们进行了大规模的模拟,以评估该程序对基因座数量和每个基因座的等位基因数量、等位基因分布类型、家族分布以及群体等位基因频率的独立知识的敏感性。基因座数量和每个基因座的等位基因数量对准确性影响最大。当它们至少有8个等位基因时,即使只有4个位点,也能获得非常好的准确性。当使用偏斜家族分布的小目标样本中使用两两似然方法估计等位基因频率时,准确性降低。我们提出了一种迭代方法,部分地纠正了这个问题。全似然方法对等位基因频率估计的准确性不太敏感,但在大数据集或可用信息很少(例如,四个等位基因的四个位点)时表现不佳。
Two Markov chain Monte Carlo algorithms are proposed that allow the partitioning of individuals into full-sib groups using single-locus genetic marker data when no parental information is available. These algorithms present a method of moving through the sibship configuration space and locating the configuration that maximizes an overall score on the basis of pairwise likelihood ratios of being full-sib or unrelated or maximizes the full joint likelihood of the proposed family structure. Using these methods, up to 757 out of 759 Atlantic salmon were correctly classified into 12 full-sib families of unequal size using four microsatellite markers. Large-scale simulations were performed to assess the sensitivity of the procedures to the number of loci and number of alleles per locus, the allelic distribution type, the distribution of families, and the independent knowledge of population allelic frequencies. The number of loci and the number of alleles per locus had the most impact on accuracy. Very good accuracy can be obtained with as few as four loci when they have at least eight alleles. Accuracy decreases when using allelic frequencies estimated in small target samples with skewed family distributions with the pairwise likelihood approach. We present an iterative approach that partly corrects that problem. The full likelihood approach is less sensitive to the precision of allelic frequencies estimates but did not perform as well with the large data set or when little information was available (e.g., four loci with four alleles).
DOI: 10.1073/pnas.87.7.2496
发表时间: 1990-04
影响因子: 11.1
作者:
H. Reeve;D. Westneat;W. Noon;P. Sherman;C. Aquadro
通讯作者: H. Reeve;D. Westneat;W. Noon;P. Sherman;C. Aquadro