Pedigree reconstruction from SNP data: parentage assignment, sibship clustering and beyond.

Pedigree reconstruction from SNP data: parentage assignment, sibship clustering and beyond.
复制标题

DOI:
10.1111/1755-0998.12665
复制
发表时间:
2017-09
影响因子:
7.7
通讯作者:
Huisman J
Huisman J
中科院分区:
生物学1区
文献类型:
--
作者:
Huisman J

文献摘要

参考文献

被引文献

相似文献

关于数百或数千个单核苷酸多态性(SNP)的数据提供了关于个体之间关系的详细信息,但目前很少有工具可以将这些信息转化为多代系谱。我提出了r包sequoia,它分配父母,聚类共享未采样父母的同父异母兄弟姐妹,并将祖父母分配给同父异母兄弟姐妹。在考虑了焦点个体之间所有可能的一级、二级和三级关系的可能性以及传统的不相关选择之后,进行了评估。这种对局部似然曲面的仔细探索是在一种快速的启发式爬山算法中实现的。当以每个焦点个体的至少一个父母为条件计算可能性时,可以区分各种类别的二级亲属。基于三个不同的大谱系(N = 1000-2000),在具有现实基因分型错误率和缺失的模拟数据集上测试性能。这包括一个复杂的谱系,世代重叠,偶尔近亲繁殖和一些未知的出生年份。亲子鉴定是高度准确的,低至约100个独立的SNP(错误率<0.1%)并且快速(<1分钟),因为大多数对可以基于相反的纯合性被排除为亲子。对于完整的系谱重建,40%的父母被假定为非基因型。当使用至少200个独立SNP时,重建在有限的计算时间(通常<1 h)内导致低错误率(<0.3%)和高分配率(>99%)。在三个经验数据集,从推断的系谱估计的相关性强相关的基因组相关性。
Data on hundreds or thousands of single nucleotide polymorphisms (SNPs) provide detailed information about the relationships between individuals, but currently few tools can turn this information into a multigenerational pedigree. I present the r package sequoia, which assigns parents, clusters half‐siblings sharing an unsampled parent and assigns grandparents to half‐sibships. Assignments are made after consideration of the likelihoods of all possible first‐, second‐ and third‐degree relationships between the focal individuals, as well as the traditional alternative of being unrelated. This careful exploration of the local likelihood surface is implemented in a fast, heuristic hill‐climbing algorithm. Distinction between the various categories of second‐degree relatives is possible when likelihoods are calculated conditional on at least one parent of each focal individual. Performance was tested on simulated data sets with realistic genotyping error rate and missingness, based on three different large pedigrees (N = 1000–2000). This included a complex pedigree with overlapping generations, occasional close inbreeding and some unknown birth years. Parentage assignment was highly accurate down to about 100 independent SNPs (error rate <0.1%) and fast (<1 min) as most pairs can be excluded from being parent–offspring based on opposite homozygosity. For full pedigree reconstruction, 40% of parents were assumed nongenotyped. Reconstruction resulted in low error rates (<0.3%), high assignment rates (>99%) in limited computation time (typically <1 h) when at least 200 independent SNPs were used. In three empirical data sets, relatedness estimated from the inferred pedigree was strongly correlated to genomic relatedness.
DOI: 10.1111/j.1420-9101.2012.02626.x
发表时间: 2012-12-01
影响因子: 2.1
作者:
Stopher, K. V.;Nussey, D. H.;Pemberton, J. M.
通讯作者: Pemberton, J. M.
DOI: 10.1007/s10592-015-0709-1
发表时间: 2015-08-01
影响因子: 2.2
作者:
Taylor, Helen R.;Kardos, Marty D.;Allendorf, Fred W.
通讯作者: Allendorf, Fred W.
DOI: 10.1111/j.1469-1809.1975.tb00120.x
发表时间: 1975-01-01
影响因子: 1.9
作者:
THOMPSON, EA
通讯作者: THOMPSON, EA
DOI: 10.1086/284554
发表时间: 1986-08-01
影响因子: 2.9
作者:
MEAGHER, TR
通讯作者: MEAGHER, TR
DOI: 10.1186/1297-9686-43-34
发表时间: 2011-10-11
期刊: Genetics, selection, evolution : GSE
影响因子: --
作者:
Calus MP;Mulder HA;Bastiaansen JW
通讯作者: Bastiaansen JW