A unified stepwise regression procedure for evaluating the relative effects of polymorphisms within a gene using case/control or family data:: Application to HLA in type 1 diabetes

A unified stepwise regression procedure for evaluating the relative effects of polymorphisms within a gene using case/control or family data:: Application to HLA in type 1 diabetes
复制标题

DOI:
10.1086/338007
复制
发表时间:
2002-01-01
影响因子:
9.8
通讯作者:
Clayton, DG
Clayton, DG
中科院分区:
生物学1区
文献类型:
--
作者:
Cordell, HJ;Clayton, DG

文献摘要

被引文献

相似文献

提出了一种逐步逻辑回归程序,用于评估一个小基因区域内不同位点变异的相对重要性。通过拟合具有主效应的统计模型,而非对完整单倍型效应进行建模,我们生成了自由度较少的检验,这些检验可能对检测主要病因决定因素具有强大的效力。该方法适用于病例/对照数据或核心家庭数据,病例/对照数据通过无条件逻辑回归建模,家庭数据通过条件逻辑回归建模。当使用家庭数据时,提出了四种不同的条件设定策略来评估多个紧密连锁位点的效应。第一种策略产生的似然性等同于对匹配的病例/对照研究进行分析,其中每个患病后代与三个伪对照相匹配,而第二种策略等同于将每个患病后代与一到三个伪对照相匹配。这两种策略都要求能够推断出亲本的单倍型(即父母中存在的那些单倍型)。无法确定单倍型的家庭必须被舍弃,这可能会大大减少数据集的有效规模,特别是当考虑大量多态性不高的位点时。因此,提出了第三种策略,该策略不需要亲本单倍型的知识,这使得那些单倍型不明确的家庭能够被纳入分析。第四种也是最后一种策略是,当能够推断出亲本单倍型时使用条件设定方法2,否则使用条件设定方法3。通过使用核心家庭数据评估HLA区域的位点对1型糖尿病发展的贡献来说明这些方法。
A stepwise logistic-regression procedure is proposed for evaluation of the relative importance of variants at different sites within a small genetic region. By fitting statistical models with main effects, rather than modeling the full haplotype effects, we generate tests, with few degrees of freedom, that are likely to be powerful for detecting primary etiological determinants. The approach is applicable to either case/control or nuclear-family data, with case/control data modeled via unconditional and family data via conditional logistic regression. Four different conditioning strategies are proposed for evaluation of effects at multiple, closely linked loci when family data are used. The first strategy results in a likelihood that is equivalent to analysis of a matched case/control study with each affected offspring matched to three pseudocontrols, whereas the second strategy is equivalent to matching each affected offspring with between one and three pseudocontrols. Both of these strategies require parental phase (i.e., those haplotypes present in the parents) to be inferable. Families in which phase cannot be determined must be discarded, which can considerably reduce the effective size of a data set, particularly when large numbers of loci that are not very polymorphic are being considered. Therefore, a third strategy is proposed in which knowledge of parental phase is not required, which allows those families with ambiguous phase to be included in the analysis. The fourth and final strategy is to use conditioning method 2 when parental phase can be inferred and to use conditioning method 3 otherwise. The methods are illustrated using nuclear-family data to evaluate the contribution of loci in the HLA region to the development of type 1 diabetes.