The impact of missing and erroneous genotypes on tagging SNP selection and power of subsequent association tests

The impact of missing and erroneous genotypes on tagging SNP selection and power of subsequent association tests
复制标题

DOI:
10.1159/000092141
复制
发表时间:
2006-01-01
期刊:
影响因子:
1.8
通讯作者:
Chase, GA
Chase, GA
中科院分区:
生物学4区
文献类型:
--
作者:
Liu, WL;Zhao, W;Chase, GA

文献摘要

被引文献

相似文献

目的:单核苷酸多态性(SNP)作为定位疾病易感基因的有效标记,但目前的基因分型技术不足以在一个典型的连锁/关联研究中对所有可用的SNP标记进行基因分型。最近已经支付了很多注意力的方法来选择最小的信息子集的SNP在确定单倍型,但一直很少调查的影响,缺失或错误的基因型对这些SNP选择算法的性能和随后的关联测试使用选定的标签SNP。本研究的目的是探讨缺失基因型或基因分型错误对标签SNP选择以及随后使用所选标签SNP进行单标记和单倍型关联检验的影响。研究方法:通过两组模拟,我们评估了三种标记SNP选择程序在存在缺失或错误基因型的情况下的性能:克莱顿的基于多样性的程序htstep、卡尔森的基于连锁不平衡(LID)的程序IdSelect和斯特拉姆的基于决定系数的程序tagsnp.exe。结果如下:当随机选择的已知基因座被重新标记为“缺失”时,我们发现通过所有三种算法选择的标记SNP的平均数量变化很小,并且使用所选择的标记SNP的随后的单个标记和单体型关联测试的功率保持接近这些测试在缺失基因型的情况下的功率。当引入随机基因分型错误时,我们发现所有三种算法选择的标记SNP的平均数量增加。在根据CYP 19区域单倍型频率模拟的数据集中,Stram程序比Carlson和克莱顿程序有更大的增加。在聚结模型下模拟的数据集中,Carlson的程序具有最大的增加,而克莱顿的程序具有最小的增加。在两组模拟中,由于存在基因分型错误,所有三个程序的单倍型检验的功效迅速下降,但单标记检验的功效没有太大降低。结论:缺失的基因型似乎对标记SNP选择和随后的单标记和单倍型关联测试没有太大影响。相反,基因分型错误可能对标记SNP选择和单倍型测试产生严重影响,但对单个标记测试没有影响。版权所有(c)2006 S. Karger AG,巴塞尔。
Objective: Single nucleotide polymorphisms (SNPs) serve as effective markers for localizing disease susceptibility genes, but current genotyping technologies are inadequate for genotyping all available SNP markers in a typical linkage/association study. Much attention has recently been paid to methods for selecting the minimal informative subset of SNPs in identifying haplotypes, but there has been little investigation of the effect of missing or erroneous genotypes on the performance of these SNP selection algorithms and subsequent association tests using the selected tagging SNPs. The purpose of this study is to explore the effect of missing genotype or genotyping error on tagging SNP selection and subsequent single marker and haplotype association tests using the selected tagging SNPs. Methods: Through two sets of simulations, we evaluated the performance of three tagging SNP selection programs in the presence of missing or erroneous genotypes: Clayton's diversity based program htstep, Carlson's linkage disequilibrium (LID) based program IdSelect, and Stram's coefficient of determination based program tagsnp.exe. Results: When randomly selected known loci were relabeled as 'missing', we found that the average number of tagging SNPs selected by all three algorithms changed very little and the power of subsequent single marker and haplotype association tests using the selected tagging SNPs remained close to the power of these tests in the absence of missing genotype. When random genotyping errors were introduced, we found that the average number of tagging SNPs selected by all three algorithms increased. In data sets simulated according to the haplotype frequecies in the CYP19 region, Stram's program had larger increase than Carlson's and Clayton's programs. In data sets simulated under the coalescent model, Carlson's program had the largest increase and Clayton's program had the smallest increase. In both sets of simulations, with the presence of genotyping errors, the power of the haplotypetests from all three programs decreased quickly, but there was not much reduction in power of the single marker tests. Conclusions: Missing genotypes do not seem to have much impact on tagging SNP selection and subsequent single marker and haplotype association tests. In contrast, genotyping errors could have severe impact on tagging SNP selection and haplotype tests, but not on single marker tests. Copyright (c) 2006 S. Karger AG, Basel.