Maximum-likelihood estimation of haplotype frequencies in nuclear families

Maximum-likelihood estimation of haplotype frequencies in nuclear families
复制标题

DOI:
10.1002/gepi.10323
复制
发表时间:
2004-07-01
影响因子:
2.1
通讯作者:
Knapp, M
Knapp, M
中科院分区:
医学4区
文献类型:
--
作者:
Becker, T;Knapp, M

文献摘要

被引文献

相似文献

单倍型分析在疾病基因关联精细定位中的重要性在过去几年中稳步增长。由于没有大规模确定单倍型的实验方法,因此必须从统计上推断阶段。对于单个基因型数据,存在几种重建技术和许多用于单倍型频率估计的期望最大化(EM)算法。最近的研究表明,结合相关个体的基因型信息大大提高了单倍型频率估计的准确性。因此,我们实现了一个用C语言编写的高度灵活的程序,称为FAMHAP,它通过em算法计算具有任意数量儿童的一般核心家庭的单倍型频率的最大似然估计(MLEs),最多可达20个snp。对于更多的位点,我们已经实现了em算法的位点迭代模式,该算法提供了多达63个SNP位点的可靠近似,或者当多等位基因标记被纳入分析时更少。缺失的基因型也可以处理。该程序能够区分病例(传播给家庭中第一个受影响的孩子的单倍型)和伪对照(关于孩子的非传播单倍型)。我们在各种模拟数据集上测试了FAMHAP的性能和获得的单倍型频率的准确性。当考虑许多标记时,该实现被证明是有效的,并且使用通常的em算法获得的估计值与使用其轨迹迭代模式获得的估计值之间没有显着差异。我们从模拟中得出结论,核心家庭的单倍型频率估计和重建的准确性通常是非常可靠的,并且对于缺失的基因型具有很强的稳健性。(C) 2004 Wiley-Liss, Inc。
The importance of haplotype analysis in the context of association fine mapping of disease genes has grown steadily over the last years. Since experimental methods to determine haplotypes on a large scale are not available, phase has to be inferred statistically. For individual genotype data, several reconstruction techniques and many implementations of the expectation-maximization (EM) algorithm for haplotype frequency estimation exist. Recent research work has shown that incorporating available genotype information of related individuals largely increases the precision of haplotype frequency estimates. We, therefore, implemented a highly flexible program written in C, called FAMHAP, which calculates maximum likelihood estimates (MLEs) of haplotype frequencies from general nuclear families with an arbitrary number of children via the EM-algorithm for up to 20 SNPs. For more loci, we have implemented a locus-iterative mode of the EM-algorithm, which gives reliable approximations of the MLEs for up to 63 SNP loci, or less when multi-allelic markers are incorporated into the analysis. Missing genotypes can be handled as well. The program is able to distinguish cases (haplotypes transmitted to the first affected child of a family) from pseudo-controls (non-transmitted haplotypes with respect to the child). We tested the performance of FAMHAP and the accuracy of the obtained haplotype frequencies on a variety of simulated data sets. The implementation proved to work well when many markers were considered and no significant differences between the estimates obtained with the usual EM-algorithm and those obtained in its locus-iterative mode were observed. We conclude from the simulations that the accuracy of haplotype frequency estimation and reconstruction in nuclear families is very reliable in general and robust against missing genotypes. (C) 2004 Wiley-Liss, Inc.