Estimation of Genetic Correlation via Linkage Disequilibrium Score Regression and Genomic Restricted Maximum Likelihood

Estimation of Genetic Correlation via Linkage Disequilibrium Score Regression and Genomic Restricted Maximum Likelihood
复制标题

DOI:
10.1016/j.ajhg.2018.03.021
复制
发表时间:
2018-06-07
影响因子:
9.8
通讯作者:
Lee, S. Hong
Lee, S. Hong
中科院分区:
生物学1区
文献类型:
--
作者:
Ni, Guiyan;Moser, Gerhard;Lee, S. Hong

文献摘要

被引文献

相似文献

遗传相关性是描述复杂性状和疾病的共同遗传结构的关键种群参数。它可以用目前最先进的方法来估计,即连锁不平衡分数回归(LDSC)和基因组约束最大似然(GREML)。与GREML相比,LDSC的计算负担大大减少,这使其成为一种有吸引力的工具,尽管LDSC估计的精度(即标准误差的大小)尚未得到彻底研究。仿真结果表明,GREML算法的精度普遍高于LDSC算法。当实际样本和用于估计LD分数的参考数据之间存在遗传异质性时,LDSC的准确性进一步下降。在估计精神分裂症(SCZ)和体重指数之间的遗传相关性的实际数据分析中,我们表明,基于类似150,000个个体的GREML估计比基于类似400,000个个体的LDSC估计(来自组合元数据)具有更高的准确性。GREML基因组分割分析表明,SCZ与身高之间的遗传相关性在调节区中显著负相关,而全基因组或LDSC方法检测能力较弱。我们的结论是,应该仔细解释LDSC估计,因为组合的元数据集之间可能存在同质性的不确定性。我们建议,在对大量复杂性状进行大规模LDSC分析后,任何有趣的发现都应该跟进,在可能的情况下,使用GREML方法进行更详细的分析,即使样本量较小。
Genetic correlation is a key population parameter that describes the shared genetic architecture of complex traits and diseases. It can be estimated by current state-of-art methods, i.e., linkage disequilibrium score regression (LDSC) and genomic restricted maximum likelihood (GREML). The massively reduced computing burden of LDSC compared to GREML makes it an attractive tool, although the accuracy (i.e., magnitude of standard errors) of LDSC estimates has not been thoroughly studied. In simulation, we show that the accuracy of GREML is generally higher than that of LDSC. When there is genetic heterogeneity between the actual sample and reference data from which LD scores are estimated, the accuracy of LDSC decreases further. In real data analyses estimating the genetic correlation between schizophrenia (SCZ) and body mass index, we show that GREML estimates based on similar to 150,000 individuals give a higher accuracy than LDSC estimates based on similar to 400,000 individuals (from combinedmeta-data). A GREML genomic partitioning analysis reveals that the genetic correlation between SCZ and height is significantly negative for regulatory regions, which whole genome or LDSC approach has less power to detect. We conclude that LDSC estimates should be carefully interpreted as there can be uncertainty about homogeneity among combined meta-datasets. We suggest that any interesting findings from massive LDSC analysis for a large number of complex traits should be followed up, where possible, with more detailed analyses with GREML methods, even if sample sizes are lesser.