Improved Use of Small Reference Panels for Conditional and Joint Analysis with GWAS Summary Statistics

Improved Use of Small Reference Panels for Conditional and Joint Analysis with GWAS Summary Statistics
复制标题

DOI:
10.1534/genetics.118.300813
复制
发表时间:
2018-06-01
期刊:
影响因子:
3.3
通讯作者:
Pan, Wei
Pan, Wei
中科院分区:
生物学2区
文献类型:
--
作者:
Deng, Yangqing;Pan, Wei

文献摘要

被引文献

相似文献

由于大规模基因组数据共享的实用性和保密性问题,通常只有元或大分析的全基因组关联研究(GWAS)总结数据是公开的,而不是个人层面的数据。重新分析这些GWAS汇总数据的广泛应用已经变得越来越普遍和有用,这通常需要使用具有个体水平基因型数据的外部参考面板来推断遗传变异之间的连锁不平衡(LD)。然而,由于样本规模很小,只有数百人,就像最受欢迎的1000基因组计划欧洲样本一样,LD的估计误差不可忽略,导致在随后的GWAS汇总数据分析中经常大幅增加假阳性数量。为了缓解在一组snp的关联检验背景下的问题,我们提出了一种替代的协方差矩阵估计器,其思想类似于多重imputation。我们使用基于模拟和真实数据的数值示例来演示使用1000基因组计划参考面板的严重问题,以及我们的新方法的改进性能。
Due to issues of practicality and confidentiality of genomic data sharing on a large scale, typically only meta-or megaanalyzed genome-wide association study (GWAS) summary data, not individual-level data, are publicly available. Reanalyses of such GWAS summary data for a wide range of applications have become more and more common and useful, which often require the use of an external reference panel with individual-level genotypic data to infer linkage disequilibrium (LD) among genetic variants. However, with a small sample size in only hundreds, as for the most popular 1000 Genomes Project European sample, estimation errors for LD are not negligible, leading to often dramatically increased numbers of false positives in subsequent analyses of GWAS summary data. To alleviate the problem in the context of association testing for a group of SNPs, we propose an alternative estimator of the covariance matrix with an idea similar to multiple imputation. We use numerical examples based on both simulated and real data to demonstrate the severe problem with the use of the 1000 Genomes Project reference panels, and the improved performance of our new approach.