Harnessing the information contained within genome-wide association studies to improve individual prediction of complex disease risk

Harnessing the information contained within genome-wide association studies to improve individual prediction of complex disease risk
复制标题

DOI:
10.1093/hmg/ddp295
复制
发表时间:
2009-09-15
影响因子:
3.5
通讯作者:
Wray, Naomi R.
Wray, Naomi R.
中科院分区:
生物学2区
文献类型:
--
作者:
Evans, David M.;Visscher, Peter M.;Wray, Naomi R.

文献摘要

被引文献

相似文献

目前遗传诊断的模式是只在已知会影响复杂疾病风险的基因座上对个体进行检测,然而现有技术可以在基因组中的数千个基因座上对个体进行基因分型。我们调查了是否可以利用全基因组关联研究的信息来提高对复杂疾病情感状态的区分。我们采用来自Wellcome Trust Case Control Consortium的全基因组数据来验证这一假设。每个疾病队列与相同的对照组一起被分成两个样本-“训练集”和“预测集”,“训练集”中鉴定了数千个可能易患疾病风险的SNP,“预测集”中评估了这些SNP的区分能力。计算预测集中每个个体的全基因组评分,包括例如个体携带的风险等位基因总数。病例对照状态对此评分进行回归,并估计受试者工作特征曲线下面积(AUC)。在大多数情况下,在全基因组评分中自由地包括SNP与更严格地选择顶级SNP相比改善了AUC,但没有基于已建立的变体进行选择。将全基因组评分添加到已知的变异信息中,只能有限地提高辨别准确性,但对双相情感障碍、冠心病和II型糖尿病最有效。我们的结论是,这个小的增加,在判别准确性是不太可能的诊断或预测效用在目前的时间。
The current paradigm within genetic diagnostics is to test individuals only at loci known to affect risk of complex disease-yet the technology exists to genotype an individual at thousands of loci across the genome. We investigated whether information from genome-wide association studies could be harnessed to improve discrimination of complex disease affection status. We employed genome-wide data from the Wellcome Trust Case Control Consortium to test this hypothesis. Each disease cohort together with the same set of controls were split into two samples-a 'Training Set', where thousands of SNPs that might predispose to disease risk were identified and a 'Prediction Set', where the discriminatory ability of these SNPs was assessed. Genome-wide scores consisting of, for example, the total number of risk alleles an individual carries was calculated for each individual in the prediction set. Case-control status was regressed on this score and the area under the receiver operator characteristic curve (AUC) estimated. In most cases, a liberal inclusion of SNPs in the genome-wide score improved AUC compared with a more stringent selection of top SNPs, but did not perform as well as selection based upon established variants. The addition of genome-wide scores to known variant information produced only a limited increase in discriminative accuracy but was most effective for bipolar disorder, coronary heart disease and type II diabetes. We conclude that this small increase in discriminative accuracy is unlikely to be of diagnostic or predictive utility at the present time.