Association analysis using next-generation sequence data from publicly available control groups: the robust variance score statistic

Association analysis using next-generation sequence data from publicly available control groups: the robust variance score statistic
复制标题

DOI:
10.1093/bioinformatics/btu196
复制
发表时间:
2014-08-01
期刊:
影响因子:
5.8
通讯作者:
Strug, Lisa J.
Strug, Lisa J.
中科院分区:
生物学3区
文献类型:
--
作者:
Derkach, Andriy;Chiang, Theodore;Strug, Lisa J.

文献摘要

被引文献

相似文献

动机:对于许多研究人员来说,使用下一代序列(NGS)数据进行足够强大的病例对照研究仍然昂贵得令人望而却步。如果可行,一个更有效的战略将是包括公开可用的顺序控制。然而,这些研究可能会被测序平台、比对、单核苷酸多态和变异调用算法、阅读深度和选择阈值的差异所混淆。假设一个人可以根据种族和其他潜在的混杂因素匹配病例和对照,并且一个人可以获得两组中排列的读数,我们调查了当比较病例和对照之间的等位基因频率时,阅读深度和选择阈值的系统性差异的影响。我们提出了一种新的基于似然的方法,即稳健方差分数(RVS),该方法用给定的观测序列数据的期望值来代替基因呼叫。结果:理论上,RVS消除了在估计小等位基因频率时的阅读深度偏差。我们还证明,使用模拟和真实的NGS数据,RVS方法可以控制I类错误,并具有与使用常见和罕见变异的真实潜在基因类型的‘金标准’分析相当的能力。
Motivation: Sufficiently powered case-control studies with next-generation sequence (NGS) data remain prohibitively expensive for many investigators. If feasible, a more efficient strategy would be to include publicly available sequenced controls. However, these studies can be confounded by differences in sequencing platform; alignment, single nucleotide polymorphism and variant calling algorithms; read depth; and selection thresholds. Assuming one can match cases and controls on the basis of ethnicity and other potential confounding factors, and one has access to the aligned reads in both groups, we investigate the effect of systematic differences in read depth and selection threshold when comparing allele frequencies between cases and controls. We propose a novel likelihood-based method, the robust variance score (RVS), that substitutes genotype calls by their expected values given observed sequence data.Results: We show theoretically that the RVS eliminates read depth bias in the estimation of minor allele frequency. We also demonstrate that, using simulated and real NGS data, the RVS method controls Type I error and has comparable power to the 'gold standard' analysis with the true underlying genotypes for both common and rare variants.