pixy: Unbiased estimation of nucleotide diversity and divergence in the presence of missing data.

pixy: Unbiased estimation of nucleotide diversity and divergence in the presence of missing data.
复制标题

DOI:
10.1111/1755-0998.13326
复制
发表时间:
2021-05
影响因子:
7.7
通讯作者:
Samuk K
Samuk K
中科院分区:
生物学1区
文献类型:
--
作者:
Korunes KL;Samuk K

文献摘要

参考文献

被引文献

相似文献

群体遗传分析通常使用汇总统计来描述遗传变异的模式,并提供对进化过程的见解。这些汇总统计量中最基本的是π和dXY,它们分别用于描述种群内和种群间的遗传多样性。在这里,我们解决了π和dXY计算中的一个普遍问题:各种类型的缺失数据产生的系统偏倚。计算π和dXY的许多流行方法都是对以变异调用格式(VCF)编码的数据进行操作,该格式通过省略不变位点来压缩遗传数据。当使用VCF计算π和dXY时,通常隐含地假设缺失的基因型(包括在VCF中未表示的位点处的那些基因型)对于参考等位基因是纯合的。在这里,我们展示了这种假设如何导致π和dXY估计值的显著向下偏倚,这与缺失数据的数量成正比。我们讨论了这个问题在群体遗传学中的普遍性和重要性,并介绍了一个用户友好的UNIX命令行实用程序,pixy,它通过一个算法来解决这个问题,该算法在缺失数据的情况下生成π和dXY的无偏估计。我们使用模拟和经验数据将pixy与现有方法进行比较,并表明pixy单独产生π和dXY的无偏估计,无论缺失数据的形式或数量如何。总之,我们的软件解决了应用群体遗传学中的一个长期存在的问题,并强调了在群体遗传分析中正确解释缺失数据的重要性。
Population genetic analyses often use summary statistics to describe patterns of genetic variation and provide insight into evolutionary processes. Among the most fundamental of these summary statistics are π and dXY, which are used to describe genetic diversity within and between populations, respectively. Here, we address a widespread issue in π and dXY calculation: systematic bias generated by missing data of various types. Many popular methods for calculating π and dXY operate on data encoded in the variant call format (VCF), which condenses genetic data by omitting invariant sites. When calculating π and dXY using a VCF, it is often implicitly assumed that missing genotypes (including those at sites not represented in the VCF) are homozygous for the reference allele. Here, we show how this assumption can result in substantial downward bias in estimates of π and dXY that is directly proportional to the amount of missing data. We discuss the pervasive nature and importance of this problem in population genetics, and introduce a user-friendly UNIX command line utility, pixy, that solves this problem via an algorithm that generates unbiased estimates of π and dXY in the face of missing data. We compare pixy to existing methods using both simulated and empirical data, and show that pixy alone produces unbiased estimates of π and dXY regardless of the form or amount of missing data. In summary, our software solves a long-standing problem in applied population genetics and highlights the importance of properly accounting for missing data in population genetic analyses.
DOI: 10.1371/journal.pcbi.1004842
发表时间: 2016-05
影响因子: 4.3
作者:
Kelleher J;Etheridge AM;McVean G
通讯作者: McVean G
DOI: 10.1186/s12859-014-0356-4
发表时间: 2014-11-25
期刊: BMC bioinformatics
影响因子: 3
作者:
Korneliussen TS;Albrechtsen A;Nielsen R
通讯作者: Nielsen R
DOI: 10.1371/journal.pone.0019379
发表时间: 2011-05-04
期刊: PloS one
影响因子: 3.7
作者:
Elshire RJ;Glaubitz JC;Sun Q;Poland JA;Kawamoto K;Buckler ES;Mitchell SE
通讯作者: Mitchell SE
DOI: 10.1038/hdy.2009.151
发表时间: 2009-12
期刊: HEREDITY
影响因子: 3.8
作者:
Noor, M. A. F.;Bennett, S. M.
通讯作者: Bennett, S. M.
DOI: 10.1093/molbev/msy224
发表时间: 2019-02-01
影响因子: 10.7
作者:
Flagel, Lex;Brandvain, Yaniv;Schrider, Daniel R.
通讯作者: Schrider, Daniel R.