Investigation of the ability of haplotype association and logistic regression to identify associated susceptibility loci

Investigation of the ability of haplotype association and logistic regression to identify associated susceptibility loci
复制标题

DOI:
10.1111/j.1469-1809.2006.00301.x
复制
发表时间:
2006-11-01
影响因子:
1.9
通讯作者:
Curtis, D.
Curtis, D.
中科院分区:
生物学4区
文献类型:
--
作者:
North, B. V.;Sham, P. C.;Curtis, D.

文献摘要

被引文献

相似文献

虽然精细间隔标记越来越多地被用于病例对照关联研究,以试图确定易感位点,但对于此类标记的最佳间隔,它们检测关联的可能能力,单标记与多标记分析的相对优点,或哪种分析方法可能是最佳的,尚不清楚。对这些问题的一些研究使用了在不同种群进化理论模型下模拟的标记。然而,HapMap项目和其他来源提供了真实的数据集,可以用来获得这些方法性能的更现实的视图。APOE周围的snp和来自两个HapMap区域的snp被用来获得多态性之间连锁不平衡(LD)关系的信息,这些LD的真实模式被用来模拟数据集,例如在病例对照研究中获得的数据集,这些snp会影响疾病的易感性。对获得的数据集进行分析,使用估计单倍型频率的异质性测试,并使用逻辑回归分析,其中仅考虑每个标记的主要影响。对假定易感位点周围的所有标记进行分析,每次使用1、2、3或4个标记。在易感位点150 kb以内的一些标记能够检测到关联。在距离小于100 kb时,与易感位点的距离与关联证据的强度之间没有相关性。当平均基因座间距为25 kb时,许多基因座无法被检测到,而当基因座间距低至2 kb时,如果基因座相对于样本量具有足够强的效应,我们可以相当肯定地说,至少有一个标记与易感基因座处于足够强的LD,从而能够检测到关联。由于基因座间距为4kb,一些易感位点在强LD中没有标记位点,这可能会削弱检测关联的能力。单倍型分析与逻辑回归相比,将每个标记的影响视为单独考虑,其性能差异不大。多标记分析有时会产生比单标记分析更显著的结果,但这种情况非常罕见。我们的结果支持这样一种观点,即如果标记是随机选择的,那么低至2 kb的间距是可取的。多标记分析有时比单标记分析更强大,所以两者都应该进行。然而,由于多标记分析很少比单标记分析更显著,因此应该强烈怀疑,当这种结果发生时,它们可能是由于基因分型错误或其他人为因素造成的。单倍型分析可能比逻辑回归更容易出现此类问题,这表明后者可能是首选方法。
While finely spaced markers are increasingly being used in case-control association studies in attempts to identify susceptibility loci, not enough is yet known as to the optimal spacing of such markers, their likely power to detect association, the relative merits of single marker versus multimarker analysis, or which methods of analysis may be optimal. Some investigations of these issues have used markers simulated under different theoretical models of population evolution. However the HapMap project and other sources provide real datasets which can be used to obtain a more realistic view of the performance of these approaches.SNPs around APOE and from two HapMap regions were used to obtain information regarding linkage disequilibrium (LD) relationships between polymorphisms, and these real patterns of LD were used to simulate datasets such as would be obtained in case-control studies were these SNPs to influence susceptibility to disease. The datasets obtained were analysed using tests for heterogeneity of estimated haplotype frequencies and using logistic regression analyses in which only main effects from each marker were considered. All markers surrounding the putative susceptibility locus were analysed, using sets of either 1, 2, 3 or 4 markers at a time.Some markers within 150 kb of the susceptibility locus were able to detect association. At distances less than 100 kb there was no correlation between the distance from the susceptibility locus and the strength of evidence for association. When the average inter-locus spacing is 25 kb many loci would not be detected, while when the spacing is as low as 2 kb one can be fairly confident that at least one marker will be in strong enough LD with the susceptibility locus to enable association to be detected, if the susceptibility locus has a strong enough effect relative to the sample size. With an inter-locus spacing of 4 kb some susceptibility loci did not have a marker locus in strong LD, potentially undermining the ability to detect association. There was little difference in the performance of haplotype-based analysis compared with logistic regression considering effects of each marker as separate. Multimarker analysis on occasion produced results which were much more highly significant than single marker analysis, but only very rarely.Our results support the view that if markers are randomly selected then a spacing as low as 2 kb is desirable. Multimarker analysis can sometimes be more powerful than single marker analysis so both should be performed. However, because it is rare for multimarker analysis to be much more highly significant than single marker analysis one should strongly suspect that when such results occur they may be due to mistakes in genotyping or through some other artefact. Haplotype analysis may be more prone to such problems than logistic regression, suggesting that the latter method might be preferred.