Evaluation of tools for identifying large copy number variations from ultra-low-coverage whole-genome sequencing data.

Evaluation of tools for identifying large copy number variations from ultra-low-coverage whole-genome sequencing data.
复制标题

DOI:
10.1186/s12864-021-07686-z
复制
发表时间:
2021-05-17
期刊:
影响因子:
4.4
通讯作者:
Elo LL
Elo LL
中科院分区:
生物学2区
文献类型:
--
作者:
Smolander J;Khan S;Singaravelu K;Kauko L;Lund RJ;Laiho A;Elo LL

文献摘要

参考文献

被引文献

相似文献

从高通量下一代全基因组测序(WGS)数据中检测拷贝数变异(CNVs)已成为近年来广泛使用的研究方法。然而,对于在各种研究和临床应用中使用的超低覆盖率(0.0005 - 0.8 x)数据,如数字核型和单细胞CNV检测,所开发的算法的适用性知之甚少。本文利用超低覆盖WGS数据,研究了六种流行的基于读取深度的CNV检测算法(bici -seq2、Canvas、CNVnator、FREEC、HMMcopy和QDNAseq)的性能。真实世界的阵列和基于核型试剂盒的验证被用作评估的基准。此外,还模拟了超低覆盖率的WGS数据,以研究这些算法识别性染色体中CNVs的能力,以及这些工具能够准确发挥作用的理论最小覆盖率。我们的研究结果表明,虽然所有的方法都能够检测到较大的CNVs,但当检测到较小的CNVs (< 2 Mbp)时,许多方法容易产生假阳性。他们在性染色体中识别CNVs的能力也存在显著的差异。总体而言,BIC-seq2被认为是统计性能最好的方法。然而,与FREEC(~ 3分钟)相比,它的显著缺点是迄今为止在所有方法中运行时间最慢(bbbb3小时),我们认为FREEC是第二好的方法。我们的对比分析表明,超低覆盖率WGS数据的CNV检测对于长度为数百万个碱基对的大拷贝数变异检测是一种非常准确的方法。这些发现促进了超低覆盖CNV检测的应用。在线版本包含补充材料,可在10.1186/s12864-021-07686-z获得。
Detection of copy number variations (CNVs) from high-throughput next-generation whole-genome sequencing (WGS) data has become a widely used research method during the recent years. However, only a little is known about the applicability of the developed algorithms to ultra-low-coverage (0.0005–0.8×) data that is used in various research and clinical applications, such as digital karyotyping and single-cell CNV detection. Here, the performance of six popular read-depth based CNV detection algorithms (BIC-seq2, Canvas, CNVnator, FREEC, HMMcopy, and QDNAseq) was studied using ultra-low-coverage WGS data. Real-world array- and karyotyping kit-based validation were used as a benchmark in the evaluation. Additionally, ultra-low-coverage WGS data was simulated to investigate the ability of the algorithms to identify CNVs in the sex chromosomes and the theoretical minimum coverage at which these tools can accurately function. Our results suggest that while all the methods were able to detect large CNVs, many methods were susceptible to producing false positives when smaller CNVs (< 2 Mbp) were detected. There was also significant variability in their ability to identify CNVs in the sex chromosomes. Overall, BIC-seq2 was found to be the best method in terms of statistical performance. However, its significant drawback was by far the slowest runtime among the methods (> 3 h) compared with FREEC (~ 3 min), which we considered the second-best method. Our comparative analysis demonstrates that CNV detection from ultra-low-coverage WGS data can be a highly accurate method for the detection of large copy number variations when their length is in millions of base pairs. These findings facilitate applications that utilize ultra-low-coverage CNV detection. The online version contains supplementary material available at 10.1186/s12864-021-07686-z.
DOI: 10.1093/bioinformatics/btq033
发表时间: 2010-03-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Quinlan AR;Hall IM
通讯作者: Hall IM
DOI: 10.1093/bioinformatics/btr670
发表时间: 2012-02-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Boeva V;Popova T;Bleakley K;Chiche P;Cappo J;Schleiermacher G;Janoueix-Lerosey I;Delattre O;Barillot E
通讯作者: Barillot E
DOI: 10.1371/journal.pone.0030377
发表时间: 2012
期刊: PloS one
影响因子: 3.7
作者:
Derrien T;Estellé J;Marco Sola S;Knowles DG;Raineri E;Guigó R;Ribeca P
通讯作者: Ribeca P
DOI: 10.1101/gr.180281.114
发表时间: 2014-11
期刊: Genome research
影响因子: 7
作者:
Ha G;Roth A;Khattra J;Ho J;Yap D;Prentice LM;Melnyk N;McPherson A;Bashashati A;Laks E;Biele J;Ding J;Le A;Rosner J;Shumansky K;Marra MA;Gilks CB;Huntsman DG;McAlpine JN;Aparicio S;Shah SP
通讯作者: Shah SP
DOI: 10.1186/s13073-016-0375-z
发表时间: 2016-11-15
期刊: GENOME MEDICINE
影响因子: 12.3
作者:
Kader, Tanjina;Goode, David L.;Gorringe, Kylie L.
通讯作者: Gorringe, Kylie L.