Estimation of dynamic SNP-heritability with Bayesian Gaussian process models.

Estimation of dynamic SNP-heritability with Bayesian Gaussian process models.
复制标题

DOI:
10.1093/bioinformatics/btaa199
复制
发表时间:
2020-06-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Sillanpää MJ
Sillanpää MJ
中科院分区:
其他
文献类型:
--
作者:
Arjas A;Hauptmann A;Sillanpää MJ

文献摘要

参考文献

被引文献

相似文献

改进的DNA技术使估计具有未知亲缘关系的远亲个体的单核苷酸多态(SNP)遗传率成为现实。对于与生长和发育相关的性状,由于过程的时间依赖性,基于纵向数据的SNP遗传力估计是有意义的。然而,到目前为止,只有很少的统计方法被用来估计动态SNP遗传力并量化其全部不确定性。我们介绍了一种完全无调谐的基于贝叶斯-高斯过程(GP)的方法来估计动态方差分量及其遗传度。对于参数估计,我们使用了一种现代的马尔可夫链蒙特卡罗方法,该方法允许完全不确定性量化。分析了几个数据集,我们的结果清楚地表明,所提出的联合估计方法(从相邻时间点‘借力’)的95%可信区间明显小于首先独立估计每个时间点的方差分量然后进行平滑的两阶段基线方法。我们使用MTG2和BLUPF90软件将该方法与随机回归模型进行了比较,定量测量表明我们的方法具有更好的性能。给出了多达1000个时间点的模拟数据和实际数据的结果。最后,我们通过对数万个个体的模拟数据验证了该方法的可扩展性。C++实现dynBGP和模拟数据可在giHub:https://github.com/aarjas/dynBGP.中获得这些程序可以在R中运行。实际数据集可在QTL档案中获得:https://phenome.jax.org/centers/QTLA.补充数据可在生物信息学在线上获得。
Improved DNA technology has made it practical to estimate single-nucleotide polymorphism (SNP)-heritability among distantly related individuals with unknown relationships. For growth- and development-related traits, it is meaningful to base SNP-heritability estimation on longitudinal data due to the time-dependency of the process. However, only few statistical methods have been developed so far for estimating dynamic SNP-heritability and quantifying its full uncertainty. We introduce a completely tuning-free Bayesian Gaussian process (GP)-based approach for estimating dynamic variance components and heritability as their function. For parameter estimation, we use a modern Markov Chain Monte Carlo method which allows full uncertainty quantification. Several datasets are analysed and our results clearly illustrate that the 95% credible intervals of the proposed joint estimation method (which ‘borrows strength’ from adjacent time points) are significantly narrower than of a two-stage baseline method that first estimates the variance components at each time point independently and then performs smoothing. We compare the method with a random regression model using MTG2 and BLUPF90 software and quantitative measures indicate superior performance of our method. Results are presented for simulated and real data with up to 1000 time points. Finally, we demonstrate scalability of the proposed method for simulated data with tens of thousands of individuals. The C++ implementation dynBGP and simulated data are available in GitHub: https://github.com/aarjas/dynBGP. The programmes can be run in R. Real datasets are available in QTL archive: https://phenome.jax.org/centers/QTLA. Supplementary data are available at Bioinformatics online.
DOI: 10.1186/1471-2156-4-s1-s21
发表时间: 2003-12-31
期刊: BMC GENETICS
影响因子: 2.9
作者:
Gee, C;Morrison, JL;Gauderman, WJ
通讯作者: Gauderman, WJ
DOI: 10.1098/rstb.2005.1669
发表时间: 2005-07-29
影响因子: 6.3
作者:
Felsenstein, J
通讯作者: Felsenstein, J
DOI: 10.1534/genetics.115.183905
发表时间: 2016-04-01
期刊: GENETICS
影响因子: 3.3
作者:
He, Liang;Sillanpaa, Mikko J.;Pitkaniemi, Janne
通讯作者: Pitkaniemi, Janne
DOI: 10.1093/bioinformatics/btw012
发表时间: 2016-05-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lee SH;van der Werf JH
通讯作者: van der Werf JH
DOI: 10.1534/g3.113.007096
发表时间: 2013-09-04
期刊: G3 (Bethesda, Md.)
影响因子: --
作者:
Kärkkäinen HP;Sillanpää MJ
通讯作者: Sillanpää MJ