Data mining approaches for genome-wide association of mood disorders.

Data mining approaches for genome-wide association of mood disorders.
复制标题

DOI:
10.1097/ypg.0b013e32834dc40d
复制
发表时间:
2012-04
影响因子:
0.9
通讯作者:
Zandi PP
Zandi PP
中科院分区:
医学4区
文献类型:
--
作者:
Pirooznia M;Seifuddin F;Judy J;Mahon PB;Bipolar Genome Study (BiGS) Consortium;Potash JB;Zandi PP

文献摘要

被引文献

相似文献

情绪障碍是主要精神疾病的高度遗传形式。随着全基因组关联研究(GWAS)的出现,阐明情绪障碍的遗传结构有望取得重大突破。然而,迄今为止,很少有易感位点被最终确定。情绪障碍的遗传病因似乎相当复杂,因此需要分析 GWAS 数据的替代方法。最近,一种捕获多个基因座等位基因影响的多基因评分方法已成功应用于精神分裂症和双相情感障碍 (BP) 的 GWAS 数据分析。然而,这种方法在处理遗传效应的复杂性时可能过于简单化。数据挖掘方法可用于分析复杂精神疾病的 GWAS 生成的高维数据。我们试图将五种数据挖掘方法(即贝叶斯网络(BN)、支持向量机(SVM)、随机森林(RF)、径向基函数网络(RBF)和逻辑回归(LR))与多基因评分方法在 BP 的 GWAS 数据分析中的性能进行比较。不同的分类方法在双相基因组研究的 GWAS 数据集(2,191 个 BP 病例和 1,434 个对照)上进行了训练,并且在 Wellcome Trust 病例对照联盟的 GWAS 数据集上测试了它们准确分类病例/对照状态的能力。通过比较接收者操作特征曲线 (AUC) 下的面积来评估测试数据集中分类器的性能。 BN 的表现是所有数据挖掘分类器中最好的,但这些分类器的表现都没有明显优于多基因评分方法。我们进一步检查了大脑中表达的基因中的 S​​NP 子集,假设这些可能与 BP 易感性最相关,但所有分类器在这组 SNP 减少的情况下表现较差。目前,所有这些方法的辨别准确性不太可能具有诊断或临床用途。需要进一步的研究来制定选择可能与疾病易感性相关的 SNP 组的策略,并确定利用其他算法来推断 SNP 组之间关系的其他数据挖掘分类器是否可能表现更好。
Mood disorders are highly heritable forms of major mental illness. A major breakthrough in elucidating the genetic architecture of mood disorders was anticipated with the advent of genome-wide association studies (GWAS). However, to date few susceptibility loci have been conclusively identified. The genetic etiology of mood disorders appears to be quite complex, and as a result, alternative approaches for analyzing GWAS data are needed. Recently, a polygenic scoring approach that captures the effects of alleles across multiple loci was successfully applied to the analysis of GWAS data in schizophrenia and bipolar disorder (BP). However, this method may be overly simplistic in its approach to the complexity of genetic effects. Data mining methods are available that may be applied to analyze the high dimensional data generated by GWAS of complex psychiatric disorders. We sought to compare the performance of five data mining methods, namely, Bayesian Networks (BN), Support Vector Machine (SVM), Random Forest (RF), Radial Basis Function network (RBF), and Logistic Regression (LR), against the polygenic scoring approach in the analysis of GWAS data on BP. The different classification methods were trained on GWAS datasets from the Bipolar Genome Study (2,191 cases with BP and 1,434 controls) and their ability to accurately classify case/control status was tested on a GWAS dataset from the Wellcome Trust Case Control Consortium. The performance of the classifiers in the test dataset was evaluated by comparing area under the receiver operating characteristic curves (AUC). BN performed the best of all the data mining classifiers, but none of these did significantly better than the polygenic score approach. We further examined a subset of SNPs in genes that are expressed in the brain, under the hypothesis that these might be most relevant to BP susceptibility, but all the classifiers performed worse with this reduced set of SNPs. The discriminative accuracy of all of these methods is unlikely to be of diagnostic or clinical utility at the present time. Further research is needed to develop strategies for selecting sets of SNPs likely to be relevant to disease susceptibility and to determine if other data mining classifiers that utilize other algorithms for inferring relationships among the sets of SNPs may perform better.