Model‐based genotype and ancestry estimation for potential hybrids with mixed‐ploidy

Model‐based genotype and ancestry estimation for potential hybrids with mixed‐ploidy
复制标题

基于模型的混合倍体潜在杂交种的基因型和祖先估计

DOI:
10.1111/1755-0998.13330
复制
发表时间:
2021
影响因子:
7.7
通讯作者:
Buerkle, C. Alex
Buerkle, C. Alex
中科院分区:
生物学1区
文献类型:
--
作者:
Shastry, Vivaswat;Adams, Paula E.;Lindtke, Dorothea;Mandeville, Elizabeth G.;Parchman, Thomas L.;Gompert, Zachariah;Buerkle, C. Alex

文献摘要

相似文献

个体之间的非随机交配可能导致遗传相似个体的空间聚类和种群分层。这种与 panmixia 的偏差在自然群体中很常见。因此,个体可以在单一群体中拥有亲缘关系,或者涉及不同群体之间的杂交。在绘制性状遗传学图谱并了解塑造个体和群体遗传变异的形成进化过程时,考虑这种混合物和结构非常重要。个体之间的分层遗传相关性通常使用源自统计模型的血统估计来量化。这些多倍体和混合倍体个体和群体模型的开发落后于二倍体模型。在这里,我们扩展并测试了一种称为熵的分层贝叶斯模型,该模型可以使用低深度序列数据来估计同源多倍体和混合倍体个体(包括个体内的性染色体和常染色体)的基因型和祖先参数。我们对模拟数据的分析说明了测序深度和基因组覆盖率之间的权衡,并发现与对基因组的较小部分进行低深度测序相比,对较大部分基因组进行低深度测序相关的误差较低。通过模拟数据以及二倍体和四倍体拟南芥种群混合分析验证了该模型具有较高的准确性和灵敏度。
Non‐random mating among individuals can lead to spatial clustering of genetically similar individuals and population stratification. This deviation from panmixia is commonly observed in natural populations. Consequently, individuals can have parentage in single populations or involving hybridization between differentiated populations. Accounting for this mixture and structure is important when mapping the genetics of traits and learning about the formative evolutionary processes that shape genetic variation among individuals and populations. Stratified genetic relatedness among individuals is commonly quantified using estimates of ancestry that are derived from a statistical model. Development of these models for polyploid and mixed‐ploidy individuals and populations has lagged behind those for diploids. Here, we extend and test a hierarchical Bayesian model, calledentropy, which can use low‐depth sequence data to estimate genotype and ancestry parameters in autopolyploid and mixed‐ploidy individuals (including sex chromosomes and autosomes within individuals). Our analysis of simulated data illustrated the trade‐off between sequencing depth and genome coverage and found lower error associated with low‐depth sequencing across a larger fraction of the genome than with high‐depth sequencing across a smaller fraction of the genome. The model has high accuracy and sensitivity as verified with simulated data and through analysis of admixture among populations of diploid and tetraploidArabidopsis arenosa.