Evaluating assembly and variant calling software for strain-resolved analysis of large DNA viruses.

Evaluating assembly and variant calling software for strain-resolved analysis of large DNA viruses.
复制标题

DOI:
10.1093/bib/bbaa123
复制
发表时间:
2021-05-20
影响因子:
9.5
通讯作者:
McHardy AC
McHardy AC
中科院分区:
生物学2区
文献类型:
--
作者:
Deng ZL;Dhingra A;Fritz A;Götting J;Münch PC;Steinbrück L;Schulz TF;Ganzenmüller T;McHardy AC

文献摘要

参考文献

被引文献

相似文献

人巨细胞病毒(HCMV)感染可导致免疫功能低下的个体和先天性感染儿童严重并发症。通过临床标本的高通量测序来表征异质性病毒群体及其进化需要准确组装单个毒株或序列变体以及合适的变体识别方法。然而,大多数方法的性能尚未评估的群体组成的低分歧病毒株与大基因组,如HCMV。在一项广泛的基准测试研究中,我们评估了15个汇编程序和6个变体调用程序,使用两种不同的库制备协议创建了10个实验室生成的基准测试数据集,以确定分析此类数据的最佳实践和挑战。大多数组装器,特别是MetaSPAdes和IVA,在回收大量菌株的一系列指标上表现良好。然而,只有一个,Savage,回收低丰度菌株,并以高度碎片化的方式。两个变异的调用者,LoFreq和VarScan 2,在所有菌株丰度中表现出色。两者都有很大一部分假阳性变体调用,这些调用在“G.G”背景下强烈富集T到G的变化。这种上下文相关的系统误差的大小与实验方案有关。我们在GNU通用公共许可证v3.0(https://github.com/hzi-bifo/Quasimodo)下提供所有基准测试数据、结果和名为QuasiModo的整个基准测试工作流程,Quasipecies Metric determination on omics,以实现对这些和其他数据的完全再现性和进一步基准测试。
Infection with human cytomegalovirus (HCMV) can cause severe complications in immunocompromised individuals and congenitally infected children. Characterizing heterogeneous viral populations and their evolution by high-throughput sequencing of clinical specimens requires the accurate assembly of individual strains or sequence variants and suitable variant calling methods. However, the performance of most methods has not been assessed for populations composed of low divergent viral strains with large genomes, such as HCMV. In an extensive benchmarking study, we evaluated 15 assemblers and 6 variant callers on 10 lab-generated benchmark data sets created with two different library preparation protocols, to identify best practices and challenges for analyzing such data. Most assemblers, especially metaSPAdes and IVA, performed well across a range of metrics in recovering abundant strains. However, only one, Savage, recovered low abundant strains and in a highly fragmented manner. Two variant callers, LoFreq and VarScan2, excelled across all strain abundances. Both shared a large fraction of false positive variant calls, which were strongly enriched in T to G changes in a ‘G.G’ context. The magnitude of this context-dependent systematic error is linked to the experimental protocol. We provide all benchmarking data, results and the entire benchmarking workflow named QuasiModo, Quasispecies Metric determination on omics, under the GNU General Public License v3.0 (https://github.com/hzi-bifo/Quasimodo), to enable full reproducibility and further benchmarking on these and other data.
DOI: 10.1073/pnas.1818130116
发表时间: 2019-03-19
影响因子: 11.1
作者:
Cudini, Juliana;Roy, Sunando;Breuer, Judith
通讯作者: Breuer, Judith
DOI: 10.1093/bioinformatics/bty919
发表时间: 2019-06-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Howison, Mark;Coetzer, Mia;Kantor, Rami
通讯作者: Kantor, Rami
DOI: 10.1093/bib/bbx079
发表时间: 2019-01-01
影响因子: 9.5
作者:
Fedonin, Gennady G.;Fantin, Yury S.;Neverov, Alexey D.
通讯作者: Neverov, Alexey D.
DOI: 10.1093/bioinformatics/btv408
发表时间: 2015-11-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Gehring JS;Fischer B;Lawrence M;Huber W
通讯作者: Huber W
DOI: 10.1186/1471-2164-15-989
发表时间: 2014-11-18
期刊: BMC genomics
影响因子: 4.4
作者:
Aguirre de Cárcer D;Angly FE;Alcamí A
通讯作者: Alcamí A