Choice of population structure informative principal components for adjustment in a case-control study

Choice of population structure informative principal components for adjustment in a case-control study
复制标题

DOI:
10.1186/1471-2156-12-64
复制
发表时间:
2011-07-19
期刊:
影响因子:
2.9
通讯作者:
Lunetta, Kathryn L.
Lunetta, Kathryn L.
中科院分区:
生物学3区
文献类型:
--
作者:
Peloso, Gina M.;Lunetta, Kathryn L.

文献摘要

被引文献

相似文献

研究背景:人口结构调整的方式有多种。目前尚不清楚最佳方法是什么,以及最佳方法是否因样本和子结构的类型而异。最简单和最直接的方法是调整捕获祖先的连续主成分(PC)。通过模拟,我们探讨了这个问题的祖先信息PC应调整的关联模型,以控制人口结构的混杂性质,同时保持最大功率。一个彻底的检查,选择PC的调整在病例对照研究中可能发生在全基因组关联study.Results的结构方案尚未报道:我们发现,当SNP和表型频率不改变的亚群,所有的选择方法提供了类似的权力和适当的I型错误的关联。当SNP没有结构化并且表型具有大结构时,则不选择PC作为协变量的选择方法通常提供最大的功效。当存在结构化SNP和非结构化表型时,在模型中包括PC的选择方法具有更大的能力。当两个SNP和表型的结构,所有的选择方法有类似的power.Conclusions:标准的做法是包括一个固定数量的PC在全基因组关联研究。根据我们的研究结果,我们得出结论,如果功率不是一个问题,那么选择相同的一组前PC用于调整所有SNP在逻辑回归是一种策略,实现适当的I型错误。然而,标准实践并非在所有情况下都是最佳的,并且为了在存在非结构化表型的情况下优化结构化SNP的功效,与测试的SNP相关的PC应包括在逻辑模型中。
Background: There are many ways to perform adjustment for population structure. It remains unclear what the optimal approach is and whether the optimal approach varies by the type of samples and substructure present. The simplest and most straightforward approach is to adjust for the continuous principal components (PCs) that capture ancestry. Through simulation, we explored the issue of which ancestry informative PCs should be adjusted for in an association model to control for the confounding nature of population structure while maintaining maximum power. A thorough examination of selecting PCs for adjustment in a case-control study across the possible structure scenarios that could occur in a genome-wide association study has not been previously reported.Results: We found that when the SNP and phenotype frequencies do not vary over the sub-populations, all methods of selection provided similar power and appropriate Type I error for association. When the SNP is not structured and the phenotype has large structure, then selection methods that do not select PCs for inclusion as covariates generally provide the most power. When there is a structured SNP and a non-structured phenotype, selection methods that include PCs in the model have greater power. When both the SNP and the phenotype are structured, all methods of selection have similar power.Conclusions: Standard practice is to include a fixed number of PCs in genome-wide association studies. Based on our findings, we conclude that if power is not a concern, then selecting the same set of top PCs for adjustment for all SNPs in logistic regression is a strategy that achieves appropriate Type I error. However, standard practice is not optimal in all scenarios and to optimize power for structured SNPs in the presence of unstructured phenotypes, PCs that are associated with the tested SNP should be included in the logistic model.