Optimized Use of Low-Depth Genotyping-by-Sequencing for Genomic Prediction Among Multi-Parental Family Pools and Single Plants in Perennial Ryegrass (Lolium perenne L.)

Optimized Use of Low-Depth Genotyping-by-Sequencing for Genomic Prediction Among Multi-Parental Family Pools and Single Plants in Perennial Ryegrass (Lolium perenne L.)
复制标题

DOI:
10.3389/fpls.2018.00369
复制
发表时间:
2018-03-21
影响因子:
5.6
通讯作者:
Janss, Luc
Janss, Luc
中科院分区:
生物学2区
文献类型:
--
作者:
Cericola, Fabio;Lenk, Ingo;Janss, Luc

文献摘要

被引文献

相似文献

黑麦草单株、双亲家系池和多亲家系池通常根据等位基因频率使用测序基因分型(GBS)分析进行基因分型。GBS检测可以在低覆盖深度进行,以降低成本。然而,减少覆盖深度会导致更高的缺失数据比例,并导致在识别每个基因座的等位基因频率时精度降低。由于后者的结果,基因组关系矩阵(GRM)将是有偏见的。GRMS中的这种偏差影响了方差估计和GBLUP用于基因组预测的准确性(GBLUP-GP)。我们推导了描述低覆盖率测序的偏差作为序列读数的二项抽样的影响的方程,并考虑了所考虑的样本的任何倍性水平。这使得我们可以在一个GRM中结合个体和池基因,将池基因类型视为多倍体基因,等于池中双亲的总倍性水平。利用模拟数据验证了三种不同黑麦草育种材料在不同覆盖深度下的GRM偏差的大小:来自单株的单株基因型、来自F-2家系的池型和来自合成品种的池型。为了更好地处理丢失的数据,我们还测试了适合分析等位基因频率基因组数据的补偿程序。利用实际数据评价了偏差校正和缺失数据填补的相对优势。我们研究了一个大型数据集,包括单株、F-2家族和合成品种,在三种GBS分析中进行了基因分型,每种分析的覆盖深度都不同,并对它们的抽穗期、冠锈病抗性和种子产量进行了评估。采用交叉验证的方法检验了GBLUP方法的准确性,证明了不同育种材料间预测的可行性。与构建GRM的标准方法相比,经偏差校正的GRM被证明提高了预测精度。在我们测试的补偿方法中,随机森林方法产生的预测准确率最高。这两种方法的结合导致了预测能力的显著提高(最高可达0.09)。跨个体和跨池预测的可能性为改进黑麦草育种计划提供了新的机会。
Ryegrass single plants, bi-parental family pools, and multi-parental family pools are often genotyped, based on allele-frequencies using genotyping-by-sequencing (GBS) assays. GBS assays can be performed at low-coverage depth to reduce costs. However, reducing the coverage depth leads to a higher proportion of missing data, and leads to a reduction in accuracy when identifying the allele-frequency at each locus. As a consequence of the latter, genomic relationship matrices (GRMs) will be biased. This bias in GRMs affects variance estimates and the accuracy of GBLUP for genomic prediction (GBLUP-GP). We derived equations that describe the bias from low-coverage sequencing as an effect of binomial sampling of sequence reads, and allowed for any ploidy level of the sample considered. This allowed us to combine individual and pool genotypes in one GRM, treating pool-genotypes as a polyploid genotype, equal to the total ploidy-level of the parents of the pool. Using simulated data, we verified the magnitude of the GRM bias at different coverage depths for three different kinds of ryegrass breeding material: individual genotypes from single plants, pool-genotypes from F-2 families, and pool-genotypes from synthetic varieties. To better handle missing data, we also tested imputation procedures, which are suited for analyzing allele-frequency genomic data. The relative advantages of the bias-correction and the imputation of missing data were evaluated using real data. We examined a large dataset, including single plants, F-2 families, and synthetic varieties genotyped in three GBS assays, each with a different coverage depth, and evaluated them for heading date, crown rust resistance, and seed yield. Cross validations were used to test the accuracy using GBLUP approaches, demonstrating the feasibility of predicting among different breeding material. Bias-corrected GRMs proved to increase predictive accuracies when compared with standard approaches to construct GRMs. Among the imputation methods we tested, the random forest method yielded the highest predictive accuracy. The combinations of these two methods resulted in a meaningful increase of predictive ability (up to 0.09). The possibility of predicting across individuals and pools provides new opportunities for improving ryegrass breeding schemes.