Sequence analysis using logic regression

Sequence analysis using logic regression
复制标题

DOI:
10.1002/gepi.2001.21.s1.s626
复制
发表时间:
2001-01-01
影响因子:
2.1
通讯作者:
Hsu, L
Hsu, L
中科院分区:
医学4区
文献类型:
--
作者:
Kooperberg, C;Ruczinski, I;Hsu, L

文献摘要

被引文献

相似文献

逻辑回归是一种新的自适应回归方法,它试图将预测值构建为(二进制)协变量的布尔组合。在本文中,我们使用该算法来处理单核苷酸多态性(SNP)序列数据。被发现的预测因素可以解释为疾病的风险因素。使用交叉验证、排列测试和独立测试集等技术来评估这些风险因素的重要性。当数据是相关的时,这些模型选择技术仍然有效,这里使用的族数据就是这种情况。在我们对基因分析研讨会12数据的分析中,我们确定了基因I和基因6上突变的确切位置,以及与受影响状态相关的基因2上的一些突变,而没有选择任何假阳性。(C)2001年Wiley-Liss,Inc.
Logic Regression is a new adaptive regression methodology that attempts to construct predictors as Boolean combinations of (binary) covariates. In this paper we use this algorithm to deal with single-nucleotide polymorphism (SNP) sequence data. The predictors that are found are interpretable as risk factors of the disease. Significance of these risk factors is assessed using techniques like cross-validation, permutation tests, and independent test sets. These model selection techniques remain valid when data is dependent, as is the case for the family data used here. In our analysis of the Genetic Analysis Workshop 12 data we identify the exact locations of mutations on gene I and gene 6 and a number of mutations on gene 2 that are associated with the affected status, without selecting any false positives. (C) 2001 Wiley-Liss, Inc.