Identification of SNP interactions using logic regression

Identification of SNP interactions using logic regression
复制标题

DOI:
10.1093/biostatistics/kxm024
复制
发表时间:
2008-01-01
期刊:
影响因子:
2.1
通讯作者:
Ickstadt, Katja
Ickstadt, Katja
中科院分区:
数学2区
文献类型:
--
作者:
Schwender, Holger;Ickstadt, Katja

文献摘要

被引文献

相似文献

单核苷酸多态性(snp)的相互作用被认为是导致散发性乳腺癌等复杂疾病的原因。因此,与此类遗传数据有关的研究的重要目标是确定导致更高患病风险的单核苷酸多态性组合,并衡量这些相互作用的重要性。有许多基于分类方法的方法,如CART和随机森林,可以测量单个变量的重要性。但是这些方法都不能直接量化变量组合的重要性。在本文中,我们展示了如何在病例对照研究中使用逻辑回归来识别解释疾病状态的SNP相互作用,并提出了两种量化这些相互作用对分类重要性的方法。然后,这些方法一方面应用于模拟数据集,另一方面应用于GENICA研究的SNP数据,该研究致力于鉴定与散发性乳腺癌相关的遗传和基因环境相互作用。
Interactions of single nucleotide polymorphisms (SNPs) are assumed to be responsible for complex diseases such as sporadic breast cancer. Important goals of studies concerned with such genetic data are thus to identify combinations of SNPs that lead to a higher risk of developing a disease and to measure the importance of these interactions. There are many approaches based on classification methods such as CART and random forests that allow measuring the importance of single variables. But none of these methods enable the importance of combinations of variables to be quantified directly. In this paper, we show how logic regression can be employed to identify SNP interactions explanatory for the disease status in a case-control study and propose 2 measures for quantifying the importance of these interactions for classification. These approaches are then applied on the one hand to simulated data sets and on the other hand to the SNP data of the GENICA study, a study dedicated to the identification of genetic and gene environment interactions associated with sporadic breast cancer.