MixClone: a mixture model for inferring tumor subclonal populations.

MixClone: a mixture model for inferring tumor subclonal populations.
复制标题

DOI:
10.1186/1471-2164-16-s2-s1
复制
发表时间:
2015
期刊:
影响因子:
4.4
通讯作者:
Xie X
Xie X
中科院分区:
生物学2区
文献类型:
--
作者:
Li Y;Xie X

文献摘要

被引文献

相似文献

肿瘤基因组通常是高度异质性的,由来自多个亚克隆类型的基因组组成。所有亚克隆类型的完整表征是肿瘤基因组分析的基本需求。随着下一代测序技术的进步,最近已经开发出计算方法来直接从癌症基因组测序数据推断肿瘤亚克隆群体。这些方法中的大多数是基于来自体细胞点突变的序列信息,然而,这些算法的准确性关键取决于由变异识别算法返回的体细胞突变的质量,并且通常需要深度覆盖以实现合理水平的准确性。我们描述了一种新的概率混合模型,MixClone,用于直接从配对的正常肿瘤样本的全基因组测序推断亚克隆群体的细胞患病率。MixClone在统一的概率框架内整合了体细胞拷贝数改变和等位基因频率的序列信息。我们使用模拟和真实的癌症测序数据集证明了该方法的实用性,并表明它显着优于现有的推断肿瘤亚克隆群体的方法。MixClone包是用Python编写的,可以在https://github.com/uci-cbcl/MixClone上公开获得。本文提出的概率混合模型为基于癌症基因组测序数据的亚克隆分析提供了一个新的框架。通过将该方法应用于模拟和真实的癌症测序数据,我们表明,整合体细胞拷贝数改变和等位基因频率的序列信息可以显着提高推断肿瘤亚克隆群体的准确性。
Tumor genomes are often highly heterogeneous, consisting of genomes from multiple subclonal types. Complete characterization of all subclonal types is a fundamental need in tumor genome analysis. With the advancement of next-generation sequencing, computational methods have recently been developed to infer tumor subclonal populations directly from cancer genome sequencing data. Most of these methods are based on sequence information from somatic point mutations, However, the accuracy of these algorithms depends crucially on the quality of the somatic mutations returned by variant calling algorithms, and usually requires a deep coverage to achieve a reasonable level of accuracy. We describe a novel probabilistic mixture model, MixClone, for inferring the cellular prevalences of subclonal populations directly from whole genome sequencing of paired normal-tumor samples. MixClone integrates sequence information of somatic copy number alterations and allele frequencies within a unified probabilistic framework. We demonstrate the utility of the method using both simulated and real cancer sequencing datasets, and show that it significantly outperforms existing methods for inferring tumor subclonal populations. The MixClone package is written in Python and is publicly available at https://github.com/uci-cbcl/MixClone. The probabilistic mixture model proposed here provides a new framework for subclonal analysis based on cancer genome sequencing data. By applying the method to both simulated and real cancer sequencing data, we show that integrating sequence information from both somatic copy number alterations and allele frequencies can significantly improve the accuracy of inferring tumor subclonal populations.