Detection of individual ploidy levels with genotyping-by-sequencing (GBS) analysis

Detection of individual ploidy levels with genotyping-by-sequencing (GBS) analysis
复制标题

DOI:
10.1111/1755-0998.12657
复制
发表时间:
2017-11-01
影响因子:
7.7
通讯作者:
Mock, Karen E.
Mock, Karen E.
中科院分区:
生物学1区
文献类型:
--
作者:
Gompert, Zachariah;Mock, Karen E.

文献摘要

被引文献

相似文献

倍性水平有时因个体或群体而异,特别是在植物中。当存在这种变异时,准确确定细胞型可以为生态学或性状变异的研究提供信息,并且是群体遗传分析所必需的。在这里,我们提出并评估了一种基于测序基因分型(GBS)数据来区分低水平倍性变异(例如二倍体、三倍体和四倍体)的统计方法。该方法根据观察到的杂合性和数千个杂合 SNP 中包含不同等位基因的 DNA 序列的比率(即等位基因比率)来推断细胞类型。尽管该方法不需要有关倍性的先验信息,但如果有的话,可以将具有已知倍性的一组参考样品纳入分析中。我们使用模拟数据集和已知包括二倍体和三倍体个体的白杨(山杨)自然种群的 GBS 数据来探索该方法的功效和局限性。该方法能够在模拟数据集中可靠地区分二倍体、三倍体和四倍体,这对于不同水平的遗传多样性、近交和种群结构都是如此。低覆盖率(即 2x)对功效和准确性的影响很小,但在分析二倍体、同源四倍体和异源四倍体的模拟混合物时有时会受到影响。当应用于来自 aspen 的 GBS 数据时,基于所提出的方法的细胞类型分配与之前的微卫星和流式细胞术数据的细胞类型分配密切匹配。 CRAN 提供了实现所提出方法的 R 包 (gbs2ploidy)。
Ploidy levels sometimes vary among individuals or populations, particularly in plants. When such variation exists, accurate determination of cytotype can inform studies of ecology or trait variation and is required for population genetic analyses. Here, we propose and evaluate a statistical approach for distinguishing low-level ploidy variants (e.g. diploids, triploids and tetraploids) based on genotyping-by-sequencing (GBS) data. The method infers cytotypes based on observed heterozygosity and the ratio of DNA sequences containing different alleles at thousands of heterozygous SNPs (i.e. allelic ratios). Whereas the method does not require prior information on ploidy, a reference set of samples with known ploidy can be included in the analysis if it is available. We explore the power and limitations of this method using simulated data sets and GBS data from natural populations of aspen (Populus tremuloides) known to include both diploid and triploid individuals. The proposed method was able to reliably discriminate among diploids, triploids and tetraploids in simulated data sets, and this was true for different levels of genetic diversity, inbreeding and population structure. Power and accuracy were minimally affected by low coverage (i.e. 2x), but did sometimes suffer when simulated mixtures of diploids, autotetraploids and allotetraploids were analysed. Cytotype assignments based on the proposed method closely matched those from previous microsatellite and flow cytometry data when applied to GBS data from aspen. An R package (gbs2ploidy) implementing the proposed method is available from CRAN.