From disease association to risk assessment: an optimistic view from genome-wide association studies on type 1 diabetes.

From disease association to risk assessment: an optimistic view from genome-wide association studies on type 1 diabetes.
复制标题

DOI:
10.1371/journal.pgen.1000678
复制
发表时间:
2009-10
期刊:
影响因子:
4.5
通讯作者:
Hakonarson H
Hakonarson H
中科院分区:
生物学2区
文献类型:
--
作者:
Wei Z;Wang K;Qu HQ;Zhang H;Bradfield J;Kim C;Frackleton E;Hou C;Glessner JT;Chiavacci R;Stanley C;Monos D;Grant SF;Polychronakos C;Hakonarson H

文献摘要

参考文献

被引文献

相似文献

全基因组关联研究(GWAS)在识别常见和复杂疾病的疾病易感位点方面卓有成效。一个遗留的问题是,我们是否能够基于基因型数据量化个体的疾病风险,以便促进复杂疾病的个性化预防和治疗。先前的研究通常未能取得令人满意的效果,主要是因为仅使用了数量有限的已确认的易感位点。在此我们提出,采用大量标记物的复杂机器学习方法可能会提高疾病风险评估的性能。我们将支持向量机(SVM)算法应用于在Affymetrix基因分型平台上生成的1型糖尿病(T1D)的GWAS数据集,并使用数百个标记物优化了一个风险评估模型。随后我们在一个具有推算基因型的独立Illumina基因分型数据集(1008例病例和1000例对照)以及一个单独的Affymetrix基因分型数据集(1529例病例和1458例对照)上对该模型进行了测试,两个数据集的受试者工作特征曲线下面积(AUC)均约为0.84。相比之下,当SVM模型或逻辑回归模型仅限于数十个已知的易感位点时,效果不佳。我们的研究表明,通过使用考虑大量标记物之间相互作用的算法,可以实现改进的疾病风险评估。我们乐观地认为,对于那些单核苷酸多态性(SNP)阵列已经捕获了相当比例风险的疾病,基于基因型的疾病风险评估可能是可行的。 全基因组关联研究(GWAS)经常被吹捧的一个用途是,其研究结果可以促进个性化医疗的实施,在个性化医疗中,针对复杂疾病的预防和治疗干预措施可以根据个体的基因图谱进行调整。然而,最近使用全基因组SNP基因型数据进行疾病风险评估的研究通常未能取得令人满意的结果,导致人们对基因型数据用于此类目的的实用性持悲观态度。在此我们提出,对包含已确认和尚未确认的疾病易感性变异的大量标记物采用复杂的机器学习方法,可能会提高疾病风险评估的性能。我们在三个大规模的1型糖尿病数据集上测试了一种称为支持向量机(SVM)的算法,并证明了对该疾病的风险评估可以非常准确。我们的结果表明,使用全基因组数据进行个体化疾病风险评估对于某些疾病(如T1D)可能比其他疾病更成功。然而,预测准确性将取决于所研究疾病的遗传力、已知的遗传风险比例以及是否使用了正确的标记物组合和正确的算法。
Genome-wide association studies (GWAS) have been fruitful in identifying disease susceptibility loci for common and complex diseases. A remaining question is whether we can quantify individual disease risk based on genotype data, in order to facilitate personalized prevention and treatment for complex diseases. Previous studies have typically failed to achieve satisfactory performance, primarily due to the use of only a limited number of confirmed susceptibility loci. Here we propose that sophisticated machine-learning approaches with a large ensemble of markers may improve the performance of disease risk assessment. We applied a Support Vector Machine (SVM) algorithm on a GWAS dataset generated on the Affymetrix genotyping platform for type 1 diabetes (T1D) and optimized a risk assessment model with hundreds of markers. We subsequently tested this model on an independent Illumina-genotyped dataset with imputed genotypes (1,008 cases and 1,000 controls), as well as a separate Affymetrix-genotyped dataset (1,529 cases and 1,458 controls), resulting in area under ROC curve (AUC) of ∼0.84 in both datasets. In contrast, poor performance was achieved when limited to dozens of known susceptibility loci in the SVM model or logistic regression model. Our study suggests that improved disease risk assessment can be achieved by using algorithms that take into account interactions between a large ensemble of markers. We are optimistic that genotype-based disease risk assessment may be feasible for diseases where a notable proportion of the risk has already been captured by SNP arrays. An often touted utility of genome-wide association studies (GWAS) is that the resulting discoveries can facilitate implementation of personalized medicine, in which preventive and therapeutic interventions for complex diseases can be tailored to individual genetic profiles. However, recent studies using whole-genome SNP genotype data for disease risk assessment have generally failed to achieve satisfactory results, leading to a pessimistic view of the utility of genotype data for such purposes. Here we propose that sophisticated machine-learning approaches on a large ensemble of markers, which contain both confirmed and as yet unconfirmed disease susceptibility variants, may improve the performance of disease risk assessment. We tested an algorithm called Support Vector Machine (SVM) on three large-scale datasets for type 1 diabetes and demonstrated that risk assessment can be highly accurate for the disease. Our results suggest that individualized disease risk assessment using whole-genome data may be more successful for some diseases (such as T1D) than other diseases. However, the predictive accuracy will be dependent on the heritability of the disease under study, the proportion of the genetic risk that is known, and that the right set of markers and right algorithms are being used.
评估18种常见遗传变异的综合遗传变异对2型糖尿病风险的综合影响。
DOI: 10.2337/db08-0504
发表时间: 2008-11
期刊: Diabetes
影响因子: 7.7
作者:
Lango H;UK Type 2 Diabetes Genetics Consortium;Palmer CN;Morris AD;Zeggini E;Hattersley AT;McCarthy MI;Frayling TM;Weedon MN
通讯作者: Weedon MN
DOI: 10.1073/pnas.97.1.262
发表时间: 2000-01-04
影响因子: 11.1
作者:
Brown, MPS;Grundy, WN;Haussler, D
通讯作者: Haussler, D
DOI: 10.1038/ng.291
发表时间: 2009-01
期刊: Nature genetics
影响因子: 30.8
作者:
Kathiresan S;Willer CJ;Peloso GM;Demissie S;Musunuru K;Schadt EE;Kaplan L;Bennett D;Li Y;Tanaka T;Voight BF;Bonnycastle LL;Jackson AU;Crawford G;Surti A;Guiducci C;Burtt NP;Parish S;Clarke R;Zelenika D;Kubalanza KA;Morken MA;Scott LJ;Stringham HM;Galan P;Swift AJ;Kuusisto J;Bergman RN;Sundvall J;Laakso M;Ferrucci L;Scheet P;Sanna S;Uda M;Yang Q;Lunetta KL;Dupuis J;de Bakker PI;O'Donnell CJ;Chambers JC;Kooner JS;Hercberg S;Meneton P;Lakatta EG;Scuteri A;Schlessinger D;Tuomilehto J;Collins FS;Groop L;Altshuler D;Collins R;Lathrop GM;Melander O;Salomaa V;Peltonen L;Orho-Melander M;Ordovas JM;Boehnke M;Abecasis GR;Mohlke KL;Cupples LA
通讯作者: Cupples LA
DOI: 10.1126/science.1141634
发表时间: 2007-05-11
期刊: SCIENCE
影响因子: 56.9
作者:
Frayling, Timothy M.;Timpson, Nicholas J.;McCarthy, Mark I.
通讯作者: McCarthy, Mark I.
DOI: 10.1093/hmg/ddn250
发表时间: 2008-10-15
影响因子: 3.5
作者:
Janssens, A. Cecile J. W.;van Duijn, Cornelia M.
通讯作者: van Duijn, Cornelia M.