Modeling Informatively Missing Genotypes in Haplotype Analysis.

Modeling Informatively Missing Genotypes in Haplotype Analysis.
复制标题

DOI:
10.1080/03610920802696588
复制
发表时间:
2009
期刊:
Communications in statistics: theory and methods
影响因子:
--
通讯作者:
Zhao H
Zhao H
中科院分区:
其他
文献类型:
--
作者:
Liu N;Bucala R;Zhao H

文献摘要

被引文献

相似文献

在实际的遗传学研究中,缺失型是很常见的。现有的大多数统计方法,包括那些关于单倍型分析的方法,都假设基因型是随机缺失的--即在给定的标记上,不同的基因型和不同的等位基因以相同的概率缺失。在我们以前的工作中,我们已经证明,违反这一假设可能会导致单倍型频率估计和单倍型关联分析的严重偏差。我们提出了一个通用的缺失数据模型来同时刻画一组两个或多个双等位标记上的缺失数据模式。我们证明了单倍型频率和缺失数据概率是可识别的,当且仅当在一般缺失数据模型下这些标记之间存在连锁不平衡。在这项研究中,我们将我们的工作扩展到多等位基因标记,并观察到类似的发现。对由两个标记组成的单倍型分析的仿真研究表明,我们提出的模型可以减少由于对缺失数据机制的错误假设而导致的单倍型频率估计的偏差。最后,我们通过硬皮病研究的真实数据集的应用来说明我们的方法的实用性。
It is common to have missing genotypes in practical genetic studies. The majority of the existing statistical methods, including those on haplotype analysis, assume that genotypes are missing at random—that is, at a given marker, different genotypes and different alleles are missing with the same probability. In our previous work, we have demonstrated that the violation of this assumption may lead to serious bias in haplotype frequency estimates and haplotype association analysis. We have proposed a general missing data model to simultaneously characterize missing data patterns across a set of two or more biallelic markers. We have proved that haplotype frequencies and missing data probabilities are identifiable if and only if there is linkage disequilibrium between these markers under the general missing data model. In this study, we extend our work to multi-allelic markers and observe a similar finding. Simulation studies on the analysis of haplotypes consisting of two markers illustrate that our proposed model can reduce the bias for haplotype frequency estimates due to incorrect assumptions on the missing data mechanism. Finally, we illustrate the utilities of our method through its application to a real data set from a study of scleroderma.