Clustering by phenotype and genome-wide association study in autism

Clustering by phenotype and genome-wide association study in autism
复制标题

DOI:
10.1038/s41398-020-00951-x
复制
发表时间:
2020-08-17
影响因子:
6.8
通讯作者:
Kuriyama, Shinichi
Kuriyama, Shinichi
中科院分区:
医学1区
文献类型:
--
作者:
Narita, Akira;Nagai, Masato;Kuriyama, Shinichi

文献摘要

被引文献

相似文献

自闭症谱系障碍(ASD)具有表型和遗传异质性特征。一项模拟研究表明,尝试将患有复杂疾病的患者分为更同质的亚组可能更有力地阐明隐藏的遗传性。我们根据 Simons Simplex Collection (SSC) 的表型变量,使用 k 均值算法进行聚类分析,聚类数量为 15。作为一项初步研究,我们进行了一项传统的全基因组关联研究 (GWAS),数据集包含 597 个 ASD 病例和 370 个对照。第二步,我们根据聚类结果划分病例,并在每个亚组与对照组中进行 GWAS(基于聚类的 GWAS)。我们还在复制阶段对另一个包含 712 个先证者和 354 个对照的 SSC 数据集进行了基于集群的 GWAS。在以传统 GWAS 设计进行的初步研究中,我们没有观察到显着的关联。在基于聚类的GWAS的第二步中,我们鉴定了65个染色体位点,其中包括位于21个基因的30个基因内位点和35个满足P < 5.0 x 10(-8)阈值的基因间位点。其中一些位点位于先前报道的 ASD 候选基因内或附近:CDH5、CNTN5、CNTNAP5、DNAH17、DPP10、DSCAM、FOXK1、GABBR2、GRIN2A5、ITPR1、NTM、SDK1、SNCA 和 SRRM4。在这 65 个显着染色体位点中,位于 SRRM4 基因内的 rs11064685 在复制队列中的病例与对照中具有显着不同的分布。这些发现表明聚类可以成功识别具有相对同质疾病病因的亚组。需要在更大的队列中进行进一步的集群验证和复制研究。
Autism spectrum disorder (ASD) has phenotypically and genetically heterogeneous characteristics. A simulation study demonstrated that attempts to categorize patients with a complex disease into more homogeneous subgroups could have more power to elucidate hidden heritability. We conducted cluster analyses using the k-means algorithm with a cluster number of 15 based on phenotypic variables from the Simons Simplex Collection (SSC). As a preliminary study, we conducted a conventional genome-wide association study (GWAS) with a data set of 597 ASD cases and 370 controls. In the second step, we divided cases based on the clustering results and conducted GWAS in each of the subgroups vs controls (cluster-based GWAS). We also conducted cluster-based GWAS on another SSC data set of 712 probands and 354 controls in the replication stage. In the preliminary study, which was conducted in conventional GWAS design, we observed no significant associations. In the second step of cluster-based GWASs, we identified 65 chromosomal loci, which included 30 intragenic loci located in 21 genes and 35 intergenic loci that satisfied the threshold ofP < 5.0 x 10(-8). Some of these loci were located within or near previously reported candidate genes for ASD:CDH5,CNTN5, CNTNAP5, DNAH17, DPP10, DSCAM,FOXK1,GABBR2, GRIN2A5,ITPR1, NTM, SDK1, SNCA, andSRRM4. Of these 65 significant chromosomal loci, rs11064685 located within theSRRM4gene had a significantly different distribution in the cases vs controls in the replication cohort. These findings suggest that clustering may successfully identify subgroups with relatively homogeneous disease etiologies. Further cluster validation and replication studies are warranted in larger cohorts.