Assessing the performance of methods for copy number aberration detection from single-cell DNA sequencing data

Assessing the performance of methods for copy number aberration detection from single-cell DNA sequencing data
复制标题

DOI:
10.1371/journal.pcbi.1008012
复制
发表时间:
2020-07-01
影响因子:
4.3
通讯作者:
Nakhleh, Luay
Nakhleh, Luay
中科院分区:
生物学2区
文献类型:
--
作者:
Mallory, Xian F.;Edrisi, Mohammadamin;Nakhleh, Luay

文献摘要

被引文献

相似文献

单细胞DNA测序技术使癌症突变及其演变轨迹的研究成为可能。体细胞拷贝数畸变(CNA)与多种癌症的发生和发展有关。已经专门针对单细胞DNA测序数据开发了一系列CNA检测方法,或对其进行了调整。了解这些方法的优势和局限性对于从单细胞DNA测序数据中获得准确的拷贝数分布非常重要。我们对三种广泛使用的方法-Ginkgo、HMMcopy和CopyNumber-在模拟和真实的数据集上进行了基准测试。为了促进这一点,我们开发了一种新的模拟器的单细胞基因组进化的CNA的存在。此外,为了评估在地面真相未知的经验数据上的性能,我们引入了一种基于概率的措施来识别潜在的错误推断。虽然单细胞DNA测序对于阐明和理解CNA非常有希望,但我们的研究结果表明,即使是最好的现有方法也不超过80%的准确性。需要显著提高这三种方法的准确性的新方法。此外,随着大数据集的生成,这些方法必须在计算上高效。
Single-cell DNA sequencing technologies are enabling the study of mutations and their evolutionary trajectories in cancer. Somatic copy number aberrations (CNAs) have been implicated in the development and progression of various types of cancer. A wide array of methods for CNA detection has been either developed specifically for or adapted to single-cell DNA sequencing data. Understanding the strengths and limitations that are unique to each of these methods is very important for obtaining accurate copy number profiles from single-cell DNA sequencing data. We benchmarked three widely used methods-Ginkgo, HMMcopy, and CopyNumber-on simulated as well as real datasets. To facilitate this, we developed a novel simulator of single-cell genome evolution in the presence of CNAs. Furthermore, to assess performance on empirical data where the ground truth is unknown, we introduce a phylogeny-based measure for identifying potentially erroneous inferences. While single-cell DNA sequencing is very promising for elucidating and understanding CNAs, our findings show that even the best existing method does not exceed 80% accuracy. New methods that significantly improve upon the accuracy of these three methods are needed. Furthermore, with the large datasets being generated, the methods must be computationally efficient.