Assessment and management of single nucleotide polymorphism genotype errors in genetic association analysis.

Assessment and management of single nucleotide polymorphism genotype errors in genetic association analysis.
复制标题

遗传关联分析中单核苷酸多态性基因型错误的评估和管理。

DOI:
10.1142/9789814447362_0003
复制
发表时间:
2001
影响因子:
--
通讯作者:
Ott,J
Ott,J
中科院分区:
--
文献类型:
--
作者:
Gordon,D;Ott,J

文献摘要

相似文献

单核苷酸多态性(SNP)可以用于病例对照设计,以测试标记(SNP)和疾病之间的关联。然而,这样的设计通常假设基因型数据报告没有错误。我们提出了一种方法,简化的随机模型方法(RPM),允许在病例对照设计中的错误,相比,全随机模型方法(FPM),假设数据是无误的。应用于2 × 2列联表的Pearson χ 2检验统计量被考虑。此外,我们提供了一种可能性的方法来估计错误率使用SNP基因型数据在CEPH家系。我们测试我们的方法(RPM)对标准方法(FPM)使用模拟数据。假设所有SNP位点都有两个等位基因,编码为1和2。我们考虑了三对错误率,两种不同的样本量,以及SNP位点的两组等位基因频率。在零假设(两个群体中等位基因频率相等)和备择假设(两个群体之间等位基因频率不同)下模拟两个群体中的SNP基因型数据。模拟总数为24次;零假设下的12次模拟和备择假设下的12次模拟。显著性水平阈值为5%。对于空值情况,9/12(75%)的模拟显示在RPM下I类错误没有增加,而3/12(25%)显示略有增加(拒绝最多7%的重复)。FPM方法的I类错误率没有增加,这也可以通过分析来证明。对于替代情况(功率),与FPM方法相比,RPM方法的功率有一致的增加,并且所考虑的模拟的平均增加为0.02。当样本量较大时,RPM和FPM方法之间的功效几乎没有差异。此外,RPM方法提供了一致的更准确的等位基因频率估计为不同的人群。我们的似然方法估计错误率与CEPH家系提供了良好的估计平均。真实错误率和我们的平均估计错误率之间的最大差异是0.006。然而,估计值存在相当大的变异性,这表明需要进行多次实验或更多的CEPH家系。研究人员可以使用本文中提出的方法来(1)估计自动基因分型过程的错误率,(2)允许关联分析中的此类错误,从而当存在错误时增加检测病例和对照群体中等位基因频率之间差异的能力。
Single nucleotide polymorphisms (SNP) may be used in case-control designs to test for association between a marker (the SNP) and a disease. However, such designs usually assume that the genotype data are reported without error. We propose a method, the reduced penetrance model method (RPM) that allows for errors in a case-control design, as compared to the full penetrance model method (FPM), that assumes data are errorless. Pearson's χ2applied to a 2 × 2 contingency table is the test statistic considered. Additionally, we provide a likelihood method to estimate error rates using SNP genotype data in CEPH pedigrees. We test our method (RPM) against the standard method (FPM) using simulated data. All SNP loci are assumed to have two alleles, coded 1 and 2.We consider three pairs of error rates, two different sample sizes, and two sets of allele frequencies for the SNP locus. SNP genotype data in two populations are simulated under a null hypothesis (allele frequencies equal in both populations) and under an alternative hypothesis (allele frequencies differ between two populations). The total number of simulations is 24; 12 simulations under the null hypothesis, and 12 simulations under the alternative. The significance level threshold is 5%.For the null case, 9/12 (75%) of the simulations show no increase in type I error under RPM, while 3/12 (25%) show a slight increase (rejecting the null for at most 7% of the replicates). There is no increase in the type I error rate for FPM method, which can also be shown analytically. For the alternative case (power), there is a consistent increase in power for the RPM method as compared to FPM method, and average increase of 0.02 for the simulations considered. When sample sizes are large there is virtually no difference in power between RPM and FPM methods. Also, the RPM method provides consistently more accurate allele frequency estimates for the various populations.Our likelihood method to estimate error rates with CEPH pedigrees provides good estimates on average. The largest difference between a true error rate and our average estimated error rate is 0.006. However, there is a fair amount of variability in the estimates, suggesting the need for multiple experiments or larger numbers of CEPH pedigrees.Researchers may use the methods presented in this paper to (1) estimate error rates for their automated genotyping process, and (2) allow for such errors in association analyses, thereby increasing power to detect differences between allele frequencies in case and control populations when errors are present.