Understanding the accuracy of statistical haplotype inference with sequence data of known phase

Understanding the accuracy of statistical haplotype inference with sequence data of known phase
复制标题

DOI:
10.1002/gepi.20185
复制
发表时间:
2007-11-01
影响因子:
2.1
通讯作者:
Hixson, James E.
Hixson, James E.
中科院分区:
医学4区
文献类型:
--
作者:
Andres, Aida M.;Clark, Andrew G.;Hixson, James E.

文献摘要

被引文献

相似文献

从无关个体的多位点基因型推断单倍型的统计方法在关联研究和群体遗传学中有重要的应用。了解影响这种推断准确性的因素是重要的,但它们的评估受到已知阶段生物数据有限的限制。我们创建了人类19号染色体单体的杂交细胞系,并在39个非洲裔美国人(AA)和欧洲裔美国人(EA)起源的个体中产生了48 kb基因组区域的单染色体完整序列。我们采用这些阶段已知的基因型和合并模拟,以评估统计单倍型重建的准确性,通过几种算法。在我们的生物学数据中,即使对于短至25-50 kb的区域,相位推断的准确性也相当低,这表明在分析重建的单倍型时需要谨慎。此外,相位推断中的估计置信度的可靠性不足以允许在后续分析中可靠地并入特定于站点的不确定性信息。我们发现,在某些混合血统(AA和EA人群)的样本中,最准确的单倍型可能是通过考虑最大的合并样本来增加样本量时获得的,尽管与这些异质性样本合并相关的假设问题。策略,以提高重建单倍型的信心,和现实的替代推断单倍型分析,进行了讨论。
Statistical methods for haplotype inference from multi-site genotypes of unrelated individuals have important application in association studies and population genetics. Understanding the factors that affect the accuracy of this inference is important, but their assessment has been restricted by the limited availability of biological data with known phase. We created hybrid cell lines monosomic for human chromosome 19 and produced single-chromosome complete sequences of a 48 kb genomic region in 39 individuals of African American (AA) and European American (EA) origin. We employ these phase-known genotypes and coalescent simulations to assess the accuracy of statistical haplotype reconstruction by several algorithms. Accuracy of phase inference was considerably low in our biological data even for regions as short as 25-50 kb, suggesting that caution is needed when analyzing reconstructed haplotypes. Moreover, the reliability of estimated confidence in phase inference is not high enough to allow for a reliable incorporation of site-specific uncertainty information in subsequent analyses. We show that, in samples of certain mixed ancestry (AA and EA populations), the most accurate haplotypes are probably obtained when increasing sample size by considering the largest, pooled sample, despite the hypothetical problems associated with pooling across those heterogeneous samples. Strategies to improve confidence in reconstructed haplotypes, and realistic alternatives to the analysis of inferred haplotypes, are discussed.