Tournaments between markers as a strategy to enhance genomic predictions

Tournaments between markers as a strategy to enhance genomic predictions
复制标题

DOI:
10.1371/journal.pone.0217283
复制
发表时间:
2019-06-24
期刊:
影响因子:
3.7
通讯作者:
Conceicao Meirelles, Sarah Laguna
Conceicao Meirelles, Sarah Laguna
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Ferreira Filho, Diogenes;de Sousa Bueno Filho, Julio Silvio;Conceicao Meirelles, Sarah Laguna

文献摘要

被引文献

相似文献

大量标记的分析对于全基因组关联研究 (GWAS) 和全基因组选择 (GWS) 至关重要。然而,有两个方法论问题限制了统计分析:高维 (p >> n) 和多重共线性。尽管有一些方法可以用于拟合高维数据的模型(例如贝叶斯套索),但在这种情况下可能出现的一个大问题是模型的预测能力对于用于拟合模型的个体应该表现良好,但对于其他个体则不应该表现良好,从而限制了模型的适用性。在调整模型来预测 GBV 之前,可以通过应用一些选择方法来减少标记数量(但保留与表型性状相关的标记)来避免这个问题。我们重新审视标记样本之间基于锦标赛的策略,其中每个样本都具有良好的估计统计特性:n>p 和低共线性。此类锦标赛是使用多元线性回归来消除标记的。该方法改编自文献中以前的作品。我们使用了模拟数据以及来自肉牛 SNP 研究的真实数据。锦标赛策略不仅可以规避 p >> n 问题,还可以最大限度地减少虚假关联。对于真实数据,当我们选择超过 20 个标记时,我们在交叉验证方案的验证组中获得了预测基因组育种值 (GBV) 与表型之间大于 0.70 的相关性;当我们选择更多标记(超过 100 个)时,相关性超过 0.90,显示了 GWAS 和 GWS 识别相关 SNP(或分离)的效率。在模拟研究中,我们得到了类似的结果。
Analysis of a large number of markers is crucial in both genome-wide association studies (GWAS) and genome-wide selection (GWS). However there are two methodological issues that restrict statistical analysis: high dimensionality (p >> n) and multicollinearity. Although there are methodologies that can be used to fit models for data with high dimensionality (eg, the Bayesian Lasso), a big problem that can occurs in this cases is that the predictive ability of the model should perform well for the individuals used to fit the model, but should not perform well for other individuals, restricting the applicability of the model. This problem can be circumvent by applying some selection methodology to reduce the number of markers (but keeping the markers associated with the phenotypic trait) before adjusting a model to predict GBVs. We revisit a tournament-based strategy between marker samples, where each sample has good statistical properties for estimation: n>p and low collinearity. Such tournaments are elaborated using multiple linear regression to eliminate markers. This method is adapted from previous works found in the literature. We used simulated data as well as real data derived from a study with SNPs in beef cattle. Tournament strategies not only circumvent the p >> n issue, but also minimize spurious associations. For real data, when we selected a few more than 20 markers, we obtained correlations greater than 0.70 between predicted Genomic Breeding Values (GBVs) and phenotypes in validation groups of a cross-validation scheme; and when we selected a larger number of markers (more than 100), the correlations exceeded 0.90, showing the efficiency in identifying relevant SNPs (or segregations) for both GWAS and GWS. In the simulation study, we obtained similar results.