Priors in Whole-Genome Regression: The Bayesian Alphabet Returns

Priors in Whole-Genome Regression: The Bayesian Alphabet Returns
复制标题

DOI:
10.1534/genetics.113.151753
复制
发表时间:
2013-07-01
期刊:
影响因子:
3.3
通讯作者:
Gianola, Daniel
Gianola, Daniel
中科院分区:
生物学2区
文献类型:
--
作者:
Gianola, Daniel

文献摘要

被引文献

相似文献

全基因组对复杂性状的预测在动植物育种中受到了极大的关注,并且正在进入人类甚至果蝇遗传学领域。术语“贝叶斯字母表”表示越来越多的字母表字母,用于表示采用的先验不同但共享相同采样模型的各种贝叶斯线性回归。我们探讨了先验分布在全基因组回归模型中的作用,以在目前基因组数据的标准情况下剖析复杂性状,其中未知参数的数量 (p) 通常超过样本大小 (n)。字母表中的成员旨在以各种方式应对这种过度参数化,但这里表明先验总是有影响力的,除非 n >> p。发生这种情况是因为参数没有被识别,因此贝叶斯学习是不完善的。由于推论并非没有先验的影响,因此应谨慎对待这些方法中关于遗传结构的主张。然而,所有这些程序都可以提供对复杂性状的合理预测,前提是通过正确进行的交叉验证来评估某些参数(“调节旋钮”)。结论是,字母表成员在表型的全基因组预测中有一席之地,但其推论价值有些可疑,至少当样本量达到 n
Whole-genome enabled prediction of complex traits has received enormous attention in animal and plant breeding and is making inroads into human and even Drosophila genetics. The term "Bayesian alphabet" denotes a growing number of letters of the alphabet used to denote various Bayesian linear regressions that differ in the priors adopted, while sharing the same sampling model. We explore the role of the prior distribution in whole-genome regression models for dissecting complex traits in what is now a standard situation with genomic data where the number of unknown parameters (p) typically exceeds sample size (n). Members of the alphabet aim to confront this overparameterization in various manners, but it is shown here that the prior is always influential, unless n >> p. This happens because parameters are not likelihood identified, so Bayesian learning is imperfect. Since inferences are not devoid of the influence of the prior, claims about genetic architecture from these methods should be taken with caution. However, all such procedures may deliver reasonable predictions of complex traits, provided that some parameters ("tuning knobs") are assessed via a properly conducted cross-validation. It is concluded that members of the alphabet have a room in whole-genome prediction of phenotypes, but have somewhat doubtful inferential value, at least when sample size is such that n