A new model calling procedure for Illumina BeadArray data.

A new model calling procedure for Illumina BeadArray data.
复制标题

DOI:
10.1186/s12863-016-0398-x
复制
发表时间:
2016-06-24
期刊:
影响因子:
2.9
通讯作者:
Li G
Li G
中科院分区:
生物学3区
文献类型:
--
作者:
Li G

文献摘要

被引文献

相似文献

准确的基因型要求高通量的Illumina数据是一个重要的步骤,以提取更多的遗传信息,为大规模的全基因组关联研究。许多流行的调用算法使用混合模型以快速有效的方式推断大量单核苷酸多态性的基因型。在实践中,混合模型主要限于推断常见SNP的基因型,其中它们的次要等位基因频率相当大。然而,准确地对罕见变异进行基因分型仍然是具有挑战性的,特别是对于一些罕见变异,其基因型的边界没有明确定义。为了进一步提高罕见变异的基因型识别准确性和质量,提出了一种新的模型识别程序,命名为M-D,用于推断Illumina BeadArray数据的基因型。在这个调用过程中,一个高斯混合模型和狄利克雷过程高斯混合模型集成来推断基因型。Illumina数据的应用表明,与其他流行的基因分型算法相比,这种新方法可以提高调用性能。
Accurate genotype calling for high throughput Illumina data is an important step to extract more genetic information for a large scale genome wide association studies. Many popular calling algorithms use mixture models to infer genotypes of a large number of single nucleotide polymorphisms in a fast and efficient way. In practice, mixture models are mostly restricted to infer genotypes for common SNPs where their minor allele frequencies are quite large. However, it is still challenging to accurately genotype rare variants, especially for some rare variants where the boundaries of their genotypes are not clearly defined. To further improve the call accuracy and the quality of genotypes on rare variants, a new model calling procedure, named M-D, is proposed to infer genotypes for the Illumina BeadArray data. In this calling procedure, a Gaussian Mixture Model and a Dirichlet Process Gaussian Mixture Model are integrated to infer genotypes. Applications to Illumina data illustrate that this new approach can improve calling performance compared to other popular genotyping algorithms.