Accuracy of haplotype frequency estimation for biallelic loci, via the expectation-maximization algorithm for unphased diploid genotype data

Accuracy of haplotype frequency estimation for biallelic loci, via the expectation-maximization algorithm for unphased diploid genotype data
复制标题

DOI:
10.1086/303069
复制
发表时间:
2000-10-01
影响因子:
9.8
通讯作者:
Schork, NJ
Schork, NJ
中科院分区:
生物学1区
文献类型:
--
作者:
Fallin, D;Schork, NJ

文献摘要

被引文献

相似文献

单倍型分析在人类疾病的遗传学研究中变得越来越普遍,因为它们能够识别可能含有疾病易感基因的独特染色体片段。单倍型的研究也被用来调查许多人口过程,如迁移和移民率,连锁不平衡强度,以及人口的相关性。不幸的是,许多单倍型分析方法需要的相位信息,可以很难从非单倍体物种的样品中获得。然而,有估计单倍型频率的策略,从非定相的二倍体基因型数据收集的样本的个人,利用期望最大化(EM)算法,以克服丢失的相位信息。与其他相位确定方法相比,必须先评估这种方法的准确性,然后才能提倡使用。在这项研究中,我们考虑和探索EM衍生的单倍型频率估计和他们的人口参数之间的误差来源,注意到这种误差大部分是由于采样误差,这是固有的所有研究,即使可以确定相位。鉴于此,我们专注于在一个样本数据集内的单倍型频率和EM衍生的单倍型频率估计所产生的估计程序之间的额外误差。我们评估了单倍型频率估计的准确性作为一个函数的一些因素,包括样本量,研究的位点数,等位基因频率,和特定位点的等位基因偏离哈代-温伯格和连锁平衡。我们指出的抽样误差和估计误差的相对影响,提请注意EM估计的显著准确性,一旦抽样误差已占。我们还建议,许多因素,可能会影响准确性,可以在一个数据集内进行经验评估的事实,可用于创建“诊断”,用户可以转向评估潜在的不准确估计。
Haplotype analyses have become increasingly common in genetic studies of human disease because of their ability to identify unique chromosomal segments likely to harbor disease-predisposing genes. The study of haplotypes is also used to investigate many population processes, such as migration and immigration rates, linkage-disequilibrium strength, and the relatedness of populations. Unfortunately, many haplotype-analysis methods require phase information that can be difficult to obtain from samples of nonhaploid species. There are, however, strategies for estimating haplotype frequencies from unphased diploid genotype data collected on a sample of individuals that make use of the expectation-maximization (EM) algorithm to overcome the missing phase information. The accuracy of such strategies, compared with other phase-determination methods, must be assessed before their use can be advocated. In this study we consider and explore sources of error between EM-derived haplotype frequency estimates and their population parameters, noting that much of this error is due to sampling error, which is inherent in all studies, even when phase can be determined. In light of this, we focus on the additional error between haplotype frequencies within a sample data set and EM-derived haplotype frequency estimates incurred by the estimation procedure. We assess the accuracy of haplotype frequency estimation as a function of a number of factors, including sample size, number of loci studied, allele frequencies, and locus-specific alellic departures from Hardy-Weinberg and linkage equilibrium. We point out the relative impacts of sampling error and estimation error, calling attention to the pronounced accuracy of EM estimates once sampling error has been accounted for. We also suggest that many factors that may influence accuracy can be assessed empirically within a data set-a fact that can be used to Create "diagnostics" that a user can turn to for assessing potential inaccuracies in estimation.