Haplotype and missing data inference in nuclear families

Haplotype and missing data inference in nuclear families
复制标题

DOI:
10.1101/gr.2204604
复制
发表时间:
2004-08-01
期刊:
影响因子:
7
通讯作者:
Cutler, DJ
Cutler, DJ
中科院分区:
生物学1区
文献类型:
--
作者:
Lin, S;Chakravarti, A;Cutler, DJ

文献摘要

被引文献

相似文献

用统计学方法从群体样本中确定连锁相位仅在高连锁不平衡(LD)区域内是准确的。然而,在遗传图谱研究中,受影响的个体,包括那些涉及病例和对照的个体,可能在人群中的低ILD区域上共享10 s至100 s数量级的同源序列。与此同时,从核心家庭推断阶段可能会受到缺失的家庭成员,缺失的基因型,以及某些基因型模式的非信息性的阻碍。在这项研究中,我们重新制定了我们以前的单倍型重建算法,及其相关的计算机程序,与来自人口样本以及他们的后代的信息相父母。在将我们的算法应用于100-kb的延伸中,根据具有人类LD典型水平的Wright-Fisher模型进行模拟,我们发现具有10%缺失数据的160个trios的相位重建在整个长度上是高度准确的(>90%)。此外,我们的算法调用估计等位基因状态的缺失数据在高精度(>95%)。最后,程序的输入容量很大,可以很容易地处理大于或等于1000条染色体中的数千个分离位点。
Determining linkage phase from Population samples with statistical methods is accurate only within regions of high linkage disequilibrium (LD). Yet, affected individuals in a genetic mapping study, including those involving cases and controls, may share sequences identical-by-descent stretching on the order of 10s to 100s of kilobases, quite possibly over regions of low ILD in the Population. At the same time, inferring phase from nuclear families may be hampered by missing family members, missing genotypes, and the noninformativity of certain genotype patterns. Ill this study, we reformulate our previous haplotype reconstruction algorithm, and its associated Computer program, to phase parents with information derived from population samples as well as from their offspring. In applications Of Our algorithm to 100-kb stretches, simulated in accordance to a Wright-Fisher model with typical levels of LD in humans, we find that phase reconstruction for 160 trios with 10% missing data is highly accurate (>90%) over the entire length. Furthermore, Our algorithm call estimate allelic status for missing data at high accuracy (>95%). Finally, the input capacity of the program is vast, easily handling thousands of segregating sites in greater than or equal to1000 chromosomes.