What SNP genotyping errors are most costly for genetic association studies?

What SNP genotyping errors are most costly for genetic association studies?
复制标题

DOI:
10.1002/gepi.10301
复制
发表时间:
2004-02-01
影响因子:
2.1
通讯作者:
Finch, SJ
Finch, SJ
中科院分区:
医学4区
文献类型:
--
作者:
Kang, SJ;Gordon, D;Finch, SJ

文献摘要

被引文献

相似文献

在进行遗传关联的病例/对照研究时,哪种基因型错误分类错误的成本最高,就保持恒定的渐近功效和显著性水平所需的增加样本量(SSN)而言?我们使用2 × 3卡方(2)独立性检验来回答单核苷酸多态性(SNP)的问题。我们的策略是在一个指定的备择假设下扩展卡方检验的渐近分布的非中心性参数来近似SSN,在误差参数中使用线性泰勒级数。我们考虑两种情况:第一种假设病例和对照中的真实基因型均符合Hardy-Weinberg平衡(HWE),第二种假设仅在对照中符合HWE。泰勒级数近似在各误差率小于2%时,相对误差小于1%。最昂贵的错误是将更常见的纯合子记录为不太常见的纯合子,在两种情况下,随着次要SNP等位基因频率接近0,成本系数无限增加。将更常见的纯合子错误分类为杂合子的成本也变得无限大,因为在两种情况下次要SNP等位基因频率都为0。对于这里建模的HWE的违反,将杂合子错误分类为不太常见的纯合子的成本变得很大,尽管有界。因此,使用具有小的次要等位基因频率的SNP需要仔细注意基因分型错误的频率,以确保符合功效规范。此外,自动基因分型的设计应尽量减少这些错误,其成本系数可能会变得无限大。
Which genotype misclassification errors are most costly, in terms of increased sample size necessary (SSN) to maintain constant asymptotic power and significance level, when performing case/control studies of genetic association? We answer this question for single-nucleotide polymorphisms (SNPs), using the 2 x 3 chi(2) test of independence. Our strategy is to expand the noncentrality parameter of the asymptotic distribution of the chi(2) test under a specified alternative hypothesis to approximate SSN, using a linear Taylor series in the error parameters. We consider two scenarios: the first assumes Hardy-Weinberg equilibrium (HWE) for the true genotypes in both cases and controls, and the second assumes HWE only in controls. The Taylor series approximation has a relative error of less than 1% when each error rate is less than 2%. The most costly error is recording the more common homozygote as the less common homozygote, with indefinitely increasing cost coefficient as minor SNP allele frequencies approach 0 in both scenarios. The cost of misclassifying the more common homozygote to the heterozygote also becomes indefinitely large as the minor SNP allele frequency goes to 0 under both scenarios. For the violation of HWE modeled here, the cost of misclassifying a heterozygote to the less common homozygote becomes large, although bounded. Therefore, the use of SNPs with a small minor allele frequency requires careful attention to the frequency of genotyping errors to ensure that power specifications are met. Furthermore, the design of automated genotyping should minimize those errors whose cost coefficients can become indefinitely large.