Comparing somatic mutation-callers: beyond Venn diagrams

Comparing somatic mutation-callers: beyond Venn diagrams
复制标题

DOI:
10.1186/1471-2105-14-189
复制
发表时间:
2013-06-10
期刊:
影响因子:
3
通讯作者:
Speed, Terence P.
Speed, Terence P.
中科院分区:
生物学4区
文献类型:
--
作者:
Kim, Su Yeon;Speed, Terence P.

文献摘要

被引文献

相似文献

背景资料:基于来自匹配的肿瘤-正常患者样本的DNA的体细胞突变识别是许多癌症基因组计划进行的关键任务之一。一个这样的大规模项目是癌症基因组图谱(TCGA),它现在正在从数百个配对的肿瘤正常DNA外显子组序列数据中定期编辑体细胞突变的目录。尽管如此,突变呼叫仍然非常具有挑战性。TCGA基准研究显示,即使是来自主要中心的相对较新的突变呼叫者也显示出实质性的差异。对突变识别者的评估或理解差异的来源并不简单,因为对于大多数肿瘤研究,基于独立的全外显子组DNA测序的验证数据不可用,只有选定的突变识别者的部分验证数据可用。结果:为了提供比较来自多个调用方的输出的准则,我们分析了来自TCGA基准研究的两组突变识别数据及其部分验证数据。探索了突变调用输出的各个方面以详细表征差异。为了评估多个调用者的性能,我们引入了四种不同程度地利用外部序列数据的方法,从具有独立的DNA-seq对,仅用于肿瘤样本的RNA-seq,仅原始外显子组-seq对,或没有those.Conclusions:我们的分析提供了可视化和理解多个调用者输出之间的差异的指导方针。此外,将四种评估方法应用于整个外显子组数据,我们说明了挑战,并强调了在评估多个调用者的性能时需要格外小心的各种情况。
Background: Somatic mutation-calling based on DNA from matched tumor-normal patient samples is one of the key tasks carried by many cancer genome projects. One such large-scale project is The Cancer Genome Atlas (TCGA), which is now routinely compiling catalogs of somatic mutations from hundreds of paired tumor-normal DNA exome-sequence data. Nonetheless, mutation calling is still very challenging. TCGA benchmark studies revealed that even relatively recent mutation callers from major centers showed substantial discrepancies. Evaluation of the mutation callers or understanding the sources of discrepancies is not straightforward, since for most tumor studies, validation data based on independent whole-exome DNA sequencing is not available, only partial validation data for a selected (ascertained) subset of sites.Results: To provide guidelines to comparing outputs from multiple callers, we have analyzed two sets of mutation-calling data from the TCGA benchmark studies and their partial validation data. Various aspects of the mutation-calling outputs were explored to characterize the discrepancies in detail. To assess the performances of multiple callers, we introduce four approaches utilizing the external sequence data to varying degrees, ranging from having independent DNA-seq pairs, RNA-seq for tumor samples only, the original exome-seq pairs only, or none of those.Conclusions: Our analyses provide guidelines to visualizing and understanding the discrepancies among the outputs from multiple callers. Furthermore, applying the four evaluation approaches to the whole exome data, we illustrate the challenges and highlight the various circumstances that require extra caution in assessing the performances of multiple callers.