Power and pitfalls of computational methods for inferring clone phylogenies and mutation orders from bulk sequencing data

Power and pitfalls of computational methods for inferring clone phylogenies and mutation orders from bulk sequencing data
复制标题

DOI:
10.1038/s41598-020-59006-2
复制
发表时间:
2020-02-26
期刊:
影响因子:
4.6
通讯作者:
Kumar, Sudhir
Kumar, Sudhir
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Miura, Sayaka;Vu, Tracy;Kumar, Sudhir

文献摘要

被引文献

相似文献

肿瘤具有广泛的遗传异质性,其形式为随时间推移以及在癌症的不同组织和区域中出现的不同克隆基因型。许多计算方法从收集自患者的多个肿瘤样品的群体批量测序数据产生克隆同源性。这些克隆遗传学用于推断肿瘤进展期间的突变顺序和克隆起源,使得选择适当的克隆去卷积方法至关重要。令人惊讶的是,这些方法在正确推断克隆同源性方面的绝对和相对准确性尚未得到一致的评估。因此,我们评估了七种计算方法的性能。重建的突变顺序和推断的克隆分组的准确性在方法之间变化很大。所有测试的方法都显示出正确鉴定肿瘤样品中存在的祖先克隆序列的有限能力。拷贝数改变的存在、转移性肿瘤演变期间肿瘤部位之间的多个接种事件的发生以及肿瘤之间癌细胞的广泛混合阻碍了所有测试方法的克隆的检测和克隆同源性的推断。总体而言,Clonetum,MACHINA和LICHeE显示出最高的整体准确性,但没有一种方法在所有模拟数据集上表现良好。因此,我们提出了选择数据分析方法的指南。
Tumors harbor extensive genetic heterogeneity in the form of distinct clone genotypes that arise over time and across different tissues and regions in cancer. Many computational methods produce clone phylogenies from population bulk sequencing data collected from multiple tumor samples from a patient. These clone phylogenies are used to infer mutation order and clone origins during tumor progression, rendering the selection of the appropriate clonal deconvolution method critical. Surprisingly, absolute and relative accuracies of these methods in correctly inferring clone phylogenies are yet to consistently assessed. Therefore, we evaluated the performance of seven computational methods. The accuracy of the reconstructed mutation order and inferred clone groupings varied extensively among methods. All the tested methods showed limited ability to identify ancestral clone sequences present in tumor samples correctly. The presence of copy number alterations, the occurrence of multiple seeding events among tumor sites during metastatic tumor evolution, and extensive intermixture of cancer cells among tumors hindered the detection of clones and the inference of clone phylogenies for all methods tested. Overall, CloneFinder, MACHINA, and LICHeE showed the highest overall accuracy, but none of the methods performed well for all simulated datasets. So, we present guidelines for selecting methods for data analysis.