Measurement error and variant-calling in deep Illumina sequencing of HIV

Measurement error and variant-calling in deep Illumina sequencing of HIV
复制标题

DOI:
10.1093/bioinformatics/bty919
复制
发表时间:
2019-06-15
期刊:
影响因子:
5.8
通讯作者:
Kantor, Rami
Kantor, Rami
中科院分区:
生物学3区
文献类型:
--
作者:
Howison, Mark;Coetzer, Mia;Kantor, Rami

文献摘要

被引文献

相似文献

下一代病毒基因组深度测序,特别是在Illumina平台上,越来越多地应用于HIV研究。然而,研究界没有标准的协议或方法来解释样品制备和测序过程中出现的测量误差。正确地调用高和低频率的变异,同时控制错误的变异是一个重要的前体下游的解释,如研究出现的艾滋病毒耐药突变,这反过来又具有临床应用,可以改善病人护理。首先,我们通过将hivmmer与真实的HIV质粒数据集上的其他变体调用管道进行比较来验证hivmmer。我们发现hivmmer实现了较低的错误变体率,并且所有方法都同意正确调用变体的频率。接下来,我们比较了使用引物ID测序的HIV质粒数据集的方法,引物ID是一种扩增子标记协议,旨在减少文库制备过程中的错误和扩增偏倚。我们表明,引物ID共识表现出更少的错误变体相比,变异调用管道,和hivmmer更接近这种低错误率相比,其他管道。来自引物ID共识的频率估计与变异体调用管道的频率估计没有显著差异。可用性和实施hivmmer可从https://github.com/kantorlab/hivmmer.Supplementary信息免费获得非商业用途补充数据可在Bioinformatics在线获得。
Motivation Next-generation deep sequencing of viral genomes, particularly on the Illumina platform, is increasingly applied in HIV research. Yet, there is no standard protocol or method used by the research community to account for measurement errors that arise during sample preparation and sequencing. Correctly calling high and low-frequency variants while controlling for erroneous variants is an important precursor to downstream interpretation, such as studying the emergence of HIV drug-resistance mutations, which in turn has clinical applications and can improve patient care.Results We developed a new variant-calling pipeline, hivmmer, for Illumina sequences from HIV viral genomes. First, we validated hivmmer by comparing it to other variant-calling pipelines on real HIV plasmid datasets. We found that hivmmer achieves a lower rate of erroneous variants, and that all methods agree on the frequency of correctly called variants. Next, we compared the methods on an HIV plasmid dataset that was sequenced using Primer ID, an amplicon-tagging protocol, which is designed to reduce errors and amplification bias during library preparation. We show that the Primer ID consensus exhibits fewer erroneous variants compared to the variant-calling pipelines, and that hivmmer more closely approaches this low error rate compared to the other pipelines. The frequency estimates from the Primer ID consensus do not differ significantly from those of the variant-calling pipelines.Availability and implementation hivmmer is freely available for non-commercial use from https://github.com/kantorlab/hivmmer.Supplementary informationSupplementary data are available at Bioinformatics online.