On the pros and cons of mega-analyses compared to meta-analyses for genome-wide association studies
On the pros and cons of mega-analyses compared to meta-analyses for genome-wide association studies
批准号:
360274005
负责人:
Professorin Dr. Iris M. Heid
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2017
资助国家:
德国
项目状态:
已结题
起止时间:
2016-12-31 至 2020-12-31
中文摘要
全基因组关联研究在识别与复杂疾病相关的遗传变异方面非常成功。为了增加样本量和统计功效,这些研究通常作为荟萃分析进行,结合了50万至3000万个基因分型和插补变异的研究特定汇总统计量。假设收集个体参与者数据(IPD)以在一个非常大的数据集中实现联合质量控制、联合插补和联合建模(大分析),而不是对汇总统计数据进行荟萃分析,将提高检测疾病位点的能力。然而,很少有工作评估什么方面的大分析与荟萃分析有助于在真实的数据场景中获得多少数据质量和疾病位点检测。由于基因型插补的高计算负担和缺乏大的全基因组IPD数据,这样的评估仍然是有限的。问题仍然是大规模分析能得到多少好处,以及是否值得付出努力。因此,我们将着手对大型IPD的大分析与研究特定汇总统计的荟萃分析的每个方面进行系统评价。我们将量化数据质量的提高和在真实的大基因组范围IPD中检测疾病位点的能力,并将来自真实的数据集的信息与模拟方法相结合,以扩展场景。我们将在三个层面上进行比较,质量控制,插补和建模。一个特别的焦点将是罕见的变异。根据我们以前的工作,我们处于一个独特的情况下,有一个大的IPD在我们的手(> 40.000人有和没有年龄相关性黄斑变性)。这种疾病和我们的数据提供了一个理想的角色模型,因为数据集由> 1200万个遗传变异组成,包括160,000个罕见的蛋白质改变,并显示出34个可检测的AMD基因座。该项目的结果将指导未来的遗传学研究,无论收集IPD数据进行大规模分析的挑战是否值得努力,并改善疾病位点检测。该项目的结果还将有助于了解如何分析大型多中心GWAS,其中所有数据都可作为IPD提供。
英文摘要
Genome-wide association studies have been highly successful in identifying genetic variants associated with complex diseases. To increase the sample size and statistical power, these studies are usually conducted as meta-analyses combining study-specific aggregated statistics for each of 0.5 to 30 million genotyped and imputed variants. It is hypothesized that collecting individual participant data (IPD) to enable joint quality control, joint imputation, and joint modelling in one very big data set (mega-analysis) instead of a meta-analysis of aggregated statistics would improve the ability to detect disease loci. However, there is little work evaluating what aspect of mega-analysis versus meta-analysis contributes to how much gain in data quality and disease loci detection in real data scenarios. Such evaluations are still limited due to the high computational burden of genotype imputing and the lack of large genome-wide IPD data. The question remains how much can be gained by mega-analysis and whether it is worth the effort. We thus will set out to conduct a systematic evaluation of each aspect of a mega-analysis of large IPD versus a meta-analysis of study-specific aggregated statistics. We will quantify the gain in the data quality and the ability to detect disease loci in a real large genome-wide IPD and will combine information from the real data set with simulation approaches to expand the scenarios. We will conduct the comparisons on three levels, the quality control, the imputation, and the modelling. One special focus will be rare variants.Based on our previous work, we are in the unique situation of having a large IPD at our hand (> 40.000 persons with and without age-related macular degeneration). This disease and our data provide an ideal role model, since the data set consists of > 12 million genetic variants including 160,000 that are rare protein-altering, and exhibits 34 detectable AMD loci. The results of this project will guide future genetic studies whether the challenge to gather IPD data for mega-analysis will be worth the effort and improve disease locus detection. The results of this project will also help understand how to analyse large multi-center GWAS where all data is available as IPD.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
On the differences between mega‐ and meta‐imputation and analysis exemplified on the genetics of age‐related macular degeneration
论megaâ和metaâcomputation之间的差异以及以年龄相关性黄斑变性的遗传学为例的分析
DOI:
10.1002/gepi.22204
发表时间:
2019
期刊:
Genetic Epidemiology
影响因子:
2.1
作者:
[Gorski M, Guenther F, Winkler TW, Weber BHF, Heid IM]
通讯作者:
Heid IM
DOI:
10.1186/s12920-020-00760-7
发表时间:
2020-08-26
期刊:
BMC MEDICAL GENOMICS
影响因子:
2.7
作者:
[Winkler, Thomas W., Grassmann, Felix, Weber, Bernhard H. F.]
通讯作者:
Weber, Bernhard H. F.
海外基金