Contaminant DNA in bacterial sequencing experiments is a major source of false genetic variability

Contaminant DNA in bacterial sequencing experiments is a major source of false genetic variability
复制标题

DOI:
10.1186/s12915-020-0748-z
复制
发表时间:
2020-03-02
期刊:
影响因子:
5.4
通讯作者:
Comas, Inaki
Comas, Inaki
中科院分区:
生物学2区
文献类型:
--
作者:
Goig, Galo A.;Blanco, Silvia;Comas, Inaki

文献摘要

被引文献

相似文献

背景污染物DNA是分子生物学和基因组库中众所周知的混杂因素。引人注目的是,全基因组测序(WGS)数据的分析工作流程通常不考虑污染可能引入的错误,这可能导致基础和临床研究中等位基因频率的错误评估。结果我们使用分类过滤器从20项不同研究的4000多份细菌样本中去除污染物读数,并对WGS中污染物DNA的程度和影响进行了全面评估。我们发现,污染是普遍存在的,并可能在变异分析中引入较大的偏差。我们发现,这些偏差可能导致数百个假阳性和假阴性SNP,即使是轻微污染的样本。如果在生物信息学分析过程中忽略污染,那么从测序数据中调查复杂生物性状的研究可能完全有偏见,并且我们证明,使用分类分类器去除污染物读数允许更准确的变体调用。我们使用真实的和模拟数据来评估和实施可靠的污染感知分析管道。结论随着测序技术作为越来越多地在研究和临床环境中采用的精密工具的巩固,我们的结果迫切需要实施污染感知分析管道。分类分类器是实现这种管道的强大工具。
Background Contaminant DNA is a well-known confounding factor in molecular biology and in genomic repositories. Strikingly, analysis workflows for whole-genome sequencing (WGS) data commonly do not account for errors potentially introduced by contamination, which could lead to the wrong assessment of allele frequency both in basic and clinical research. Results We used a taxonomic filter to remove contaminant reads from more than 4000 bacterial samples from 20 different studies and performed a comprehensive evaluation of the extent and impact of contaminant DNA in WGS. We found that contamination is pervasive and can introduce large biases in variant analysis. We showed that these biases can result in hundreds of false positive and negative SNPs, even for samples with slight contamination. Studies investigating complex biological traits from sequencing data can be completely biased if contamination is neglected during the bioinformatic analysis, and we demonstrate that removing contaminant reads with a taxonomic classifier permits more accurate variant calling. We used both real and simulated data to evaluate and implement reliable, contamination-aware analysis pipelines. Conclusion As sequencing technologies consolidate as precision tools that are increasingly adopted in the research and clinical context, our results urge for the implementation of contamination-aware analysis pipelines. Taxonomic classifiers are a powerful tool to implement such pipelines.