Confounded by sequencing depth in association studies of rare alleles.

Confounded by sequencing depth in association studies of rare alleles.
复制标题

DOI:
10.1002/gepi.20574
复制
发表时间:
2011-05
影响因子:
2.1
通讯作者:
Garner, Chad
Garner, Chad
中科院分区:
医学4区
文献类型:
--
作者:
Garner, Chad

文献摘要

参考文献

被引文献

相似文献

下一代DNA测序技术正在促进对稀有基因变异的大规模关联研究。序列读取覆盖的深度是下一代技术中的一个重要实验变量,它是从序列数据生成的基因呼叫质量的主要决定因素。当病例和对照样本单独测序或跨批次以不同比例测序时,它们在测序阅读深度上不太可能匹配,可能导致不同的基因类型错误分类,导致混淆和假阳性率增加。来自1000基因组计划试点研究3的数据被用来证明,病例和对照样本的平均测序阅读深度之间的差异可能导致罕见和罕见变异的假阳性关联,即使当两组的平均覆盖深度都超过30倍时也是如此。假阳性率的混淆和膨胀程度取决于病例组和对照组平均深度的不同程度。Logistic回归模型被用来检验病例对照状态与一组罕见和不常见的变异中累积的等位基因数量之间的关联。在Logistic回归模型中包括每个个体在不同位置的平均序列阅读深度,几乎消除了混杂效应和夸大的假阳性率。此外,在回归分析中通过对杂合子基因型呼叫的概率进行建模来考虑潜在的误差,对统计结果有相对较小但有益的影响。
Next-generation DNA sequencing technologies are facilitating large-scale association studies of rare genetic variants. The depth of the sequence read coverage is an important experimental variable in the next-generation technologies and it is a major determinant of the quality of genotype calls generated from sequence data. When case and control samples are sequenced separately or in different proportions across batches, they are unlikely to be matched on sequencing read depth and a differential misclassification of genotypes can result, causing confounding and an increased false positive rate. Data from Pilot Study 3 of the 1000 Genomes project was used to demonstrate that a difference between the mean sequencing read depth of case and control samples can result in false-positive association for rare and uncommon variants, even when the mean coverage depth exceeds 30X in both groups. The degree of the confounding and inflation in the false-positive rate depended on the extent to which the mean depth was different in the case and control groups. A logistic regression model was used to test for association between case-control status and the cumulative number of alleles in a collapsed set of rare and uncommon variants. Including each individual's mean sequence read depth across the variant sites in the logistic regression model nearly eliminated the confounding effect and the inflated false positive rate. Furthermore, accounting for the potential error by modeling the probability of the heterozygote genotype calls in the regression analysis had a relatively minor but beneficial effect on the statistical results.
DOI: 10.1002/gepi.20483
发表时间: 2010-07-01
影响因子: 2.1
作者:
Garner, Chad
通讯作者: Garner, Chad
DOI: 10.1093/bioinformatics/btq448
发表时间: 2010-10-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Zhou, Hua;Sehl, Mary E.;Lange, Kenneth
通讯作者: Lange, Kenneth
DOI: 10.1016/j.ajhg.2008.06.024
发表时间: 2008-09-12
影响因子: 9.8
作者:
Li, Bingshan;Leal, Suzanne M.
通讯作者: Leal, Suzanne M.
DOI: 10.1371/journal.pgen.1000130
发表时间: 2008-07-25
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Hoggart, Clive J.;Whittaker, John C.;De Iorio, Maria;Balding, David J.
通讯作者: Balding, David J.
DOI: 10.1038/ng.f.136
发表时间: 2008-06
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --