Measuring Genetic Differentiation from Pool-seq Data

Measuring Genetic Differentiation from Pool-seq Data
复制标题

DOI:
10.1534/genetics.118.300900
复制
发表时间:
2018-09-01
期刊:
影响因子:
3.3
通讯作者:
Vitalis, Renaud
Vitalis, Renaud
中科院分区:
生物学2区
文献类型:
--
作者:
Hivert, Valentin;Leblois, Raphael;Vitalis, Renaud

文献摘要

被引文献

相似文献

高通量测序和基因分型技术的出现使得能够比较大量标记的多态性模式。虽然对于许多非模型物种来说,从个体测序数据中表征遗传结构仍然很昂贵,但事实表明,个体 DNA 测序池 (Pool-seq) 是一种有吸引力且具有成本效益的替代方案。然而,分析 DNA 池中的序列读取计数而不是单个基因型,给得出遗传分化的正确估计带来了统计挑战。在本文中,我们基于方差分析框架,为 Pool-seq 数据提供了一种 F-ST 矩量法估计器。我们通过模拟表明,这种新的估计器是无偏的,并且优于之前提出的估计器。我们评估了估计器对模型错误指定的鲁棒性,例如测序错误和单个 DNA 对池的贡献不均匀。最后,通过重新分析已发表的多刺杜父鱼 Cottus asper 不同生态型的 Pool-seq 数据,我们展示了使用无偏 F-ST 估计量可能如何质疑从之前的分析中推断出的种群结构的解释。
The advent of high throughput sequencing and genotyping technologies enables the comparison of patterns of polymorphisms at a very large number of markers. While the characterization of genetic structure from individual sequencing data remains expensive for many nonmodel species, it has been shown that sequencing pools of individual DNAs (Pool-seq) represents an attractive and cost-effective alternative. However, analyzing sequence read counts from a DNA pool instead of individual genotypes raises statistical challenges in deriving correct estimates of genetic differentiation. In this article, we provide a method-of-moments estimator of F-ST for Pool-seq data, based on an analysis-of-variance framework. We show, by means of simulations, that this new estimator is unbiased and outperforms previously proposed estimators. We evaluate the robustness of our estimator to model misspecification, such as sequencing errors and uneven contributions of individual DNAs to the pools. Finally, by reanalyzing published Pool-seq data of different ecotypes of the prickly sculpin Cottus asper, we show how the use of an unbiased F-ST estimator may question the interpretation of population structure inferred from previous analyses.