Estimating error models for whole genome sequencing using mixtures of Dirichlet-multinomial distributions.
Estimating error models for whole genome sequencing using mixtures of Dirichlet-multinomial distributions.
复制标题
DOI:
10.1093/bioinformatics/btx133
复制
发表时间:
2017-08-01
期刊:
影响因子:
--
通讯作者:
Cartwright RA
中科院分区:
文献类型:
--
作者:
Wu SH;Schwartz RS;Winter DJ;Conrad DF;Cartwright RA
Accurate identification of genotypes is an essential part of the analysis of genomic data, including in identification of sequence polymorphisms, linking mutations with disease and determining mutation rates. Biological and technical processes that adversely affect genotyping include copy-number-variation, paralogous sequences, library preparation, sequencing error and reference-mapping biases, among others. We modeled the read depth for all data as a mixture of Dirichlet-multinomial distributions, resulting in significant improvements over previously used models. In most cases the best model was comprised of two distributions. The major-component distribution is similar to a binomial distribution with low error and low reference bias. The minor-component distribution is overdispersed with higher error and reference bias. We also found that sites fitting the minor component are enriched for copy number variants and low complexity regions, which can produce erroneous genotype calls. By removing sites that do not fit the major component, we can improve the accuracy of genotype calls. Methods and data files are available at https://github.com/CartwrightLab/WuEtAl2017/ (doi:10.5281/zenodo.256858). Supplementary data is available at Bioinformatics online.
登录
查看更多内容
影响因子:
3.7
作者:
Frith, Martin C.
通讯作者:
Frith, Martin C.
影响因子:
5.8
作者:
Li, Heng
通讯作者:
Li, Heng
影响因子:
5.8
作者:
Malhis, Nawar;Jones, Steven J. M.
通讯作者:
Jones, Steven J. M.
影响因子:
30.8
作者:
通讯作者:
--
影响因子:
14.9
作者:
Josephidou M;Lynch AG;Tavaré S
通讯作者:
Tavaré S