Statistical methods with exhaustive search in the identification of gene-gene interactions for colorectal cancer

Statistical methods with exhaustive search in the identification of gene-gene interactions for colorectal cancer
复制标题

DOI:
10.1002/gepi.22372
复制
发表时间:
2020-11-24
影响因子:
2.1
通讯作者:
Hu, Ting
Hu, Ting
中科院分区:
医学4区
文献类型:
--
作者:
Kafaie, Somayeh;Xu, Ling;Hu, Ting

文献摘要

被引文献

相似文献

虽然遗传力的加性形式主要是在遗传学中研究的,但非线性、非加性的基因-基因相互作用,即上位性,可以解释包括癌症在内的复杂人类疾病中遗传性缺失的一部分。近年来,人们引入了强大的计算方法来理解极高维全基因组数据中这些复杂人类疾病的多变量遗传因素。在本研究中,我们研究了三种强大方法的性能,即基于布尔运算的筛选和测试 (BOOST)、FastEpistasis 和基于树的上位关联图谱 (TEAM),以确定结直肠癌 (CRC) 的相互作用的遗传风险因素,用于全基因组关联研究 (GWAS)。经过基于质量控制的数据预处理后,我们将这三种算法应用于CRC GWAS数据集,并选择每种方法识别出的排名最高的100个单核苷酸多态性(SNP)对(总共251个SNP),其中74对在FastEpistasis和BOOST之间是共同的。 BOOST、FastEpistasis 和 TEAM 识别的 SNP 分别映射到 58、57 和 62 个基因。我们研究中强调的一些基因,包括 MACF1、USP49、SMAD2、SMAD3、TGFBR1 和 RHOA,已在之前的 CRC 相关研究中检测到。我们还鉴定了一些与 CRC 具有潜在生物学相关性的新基因,例如 CCDC32。此外,我们为三种方法构建了这些顶级 SNP 对的网络,网络中识别的模式表明,包括 rs2412531、rs349699 和 rs17142011 在内的一些 SNP 在我们的研究中的疾病状态分类中发挥着至关重要的作用。
Though additive forms of heritability are primarily studied in genetics, nonlinear, non-additive gene-gene interactions, that is, epistasis, could explain a portion of the missing heritability in complex human diseases including cancer. In recent years, powerful computational methods have been introduced to understand multivariable genetic factors of these complex human diseases in extremely high-dimensional genome-wide data. In this study, we investigated the performance of three powerful methods, BOolean Operation-based Screening and Testing (BOOST), FastEpistasis, and Tree-based Epistasis Association Mapping (TEAM) to identify interacting genetic risk factors of colorectal cancer (CRC) for genome-wide association studies (GWAS). After quality-control based data preprocessing, we applied these three algorithms to a CRC GWAS data set, and selected the top-ranked 100 single-nucleotide polymorphism (SNP) pairs identified by each method (251 SNPs in total), among which 74 pairs were common between FastEpistasis and BOOST. The identified SNPs by BOOST, FastEpistasis, and TEAM mapped to 58, 57, and 62 genes, respectively. Some genes highlighted by our study, including MACF1, USP49, SMAD2, SMAD3, TGFBR1, and RHOA, have been detected in previous CRC-related research. We also identified some new genes with potential biological relevance to CRC such as CCDC32. Furthermore, we constructed the network of these top SNP pairs for three methods, and the patterns identified in the networks show that some SNPs including rs2412531, rs349699, and rs17142011 play a crucial role in the classification of disease status in our study.