A BAYESIAN NONPARAMETRIC MODEL FOR INFERRING SUBCLONAL POPULATIONS FROM STRUCTURED DNA SEQUENCING DATA.

A BAYESIAN NONPARAMETRIC MODEL FOR INFERRING SUBCLONAL POPULATIONS FROM STRUCTURED DNA SEQUENCING DATA.
复制标题

从结构DNA测序数据推断亚克隆群体的贝叶斯非参数模型

DOI:
10.1214/20-aoas1434
复制
发表时间:
2021-06
期刊:
The annals of applied statistics
影响因子:
--
通讯作者:
Flaherty P
Flaherty P
中科院分区:
其他
文献类型:
--
作者:
He S;Schein A;Sarsani V;Flaherty P

文献摘要

参考文献

相似文献

在肿瘤、个体和癌症类型中都可以发现癌症的不同特征或“特征”,这些特征可以由特定的基因突变驱动。然而,单细胞和大量DNA测序数据证明,在单个肿瘤内往往存在广泛的遗传异质性。这项工作的目标是通过整合单细胞和批量测序数据,联合推断肿瘤亚群的潜在基因类型和这些亚群在单个肿瘤中的分布。在治疗时了解肿瘤的遗传成分对于个性化设计靶向治疗组合和监测治疗后可能的复发非常重要。我们提出了一种分层Dirichlet过程混合模型,该模型结合了结构化抽样安排所产生的相关性结构,并证明了该模型提高了推理质量。我们将递阶Dirichlet过程先验表示为Gamma-Poisson阶系,并利用这种表示,利用增广边际化方法,推导出一种快速的Gibbs抽样推理算法。用模拟数据进行的实验表明,对于混合计数数据的分解,我们的模型优于标准的数值和统计方法。对真实的急性淋巴细胞白血病癌症测序数据集的分析表明,我们的模型改进了最先进的生物信息学方法。对我们的模型在这个真实数据集上的结果的解释揭示了样本中的共同突变的基因座。
There are distinguishing features or “hallmarks” of cancer that are found across tumors, individuals, and types of cancer, and these hallmarks can be driven by specific genetic mutations. Yet, within a single tumor there is often extensive genetic heterogeneity as evidenced by single-cell and bulk DNA sequencing data. The goal of this work is to jointly infer the underlying genotypes of tumor subpopulations and the distribution of those subpopulations in individual tumors by integrating single-cell and bulk sequencing data. Understanding the genetic composition of the tumor at the time of treatment is important in the personalized design of targeted therapeutic combinations and monitoring for possible recurrence after treatment. We propose a hierarchical Dirichlet process mixture model that incorporates the correlation structure induced by a structured sampling arrangement and we show that this model improves the quality of inference. We develop a representation of the hierarchical Dirichlet process prior as a Gamma-Poisson hierarchy and we use this representation to derive a fast Gibbs sampling inference algorithm using the augment-and-marginalize method. Experiments with simulation data show that our model outperforms standard numerical and statistical methods for decomposing admixed count data. Analyses of real acute lymphoblastic leukemia cancer sequencing dataset shows that our model improves upon state-of-the-art bioinformatic methods. An interpretation of the results of our model on this real dataset reveals co-mutated loci across samples.
DOI: 10.1158/0008-5472.can-11-0153
发表时间: 2011-06-15
期刊: Cancer research
影响因子: 11.2
作者:
Bonavia R;Inda MM;Cavenee WK;Furnari FB
通讯作者: Furnari FB
DOI: 10.1073/pnas.0801523105
发表时间: 2008-09-02
影响因子: 11.1
作者:
Campbell, Peter J.;Pleasance, Erin D.;Stratton, Michael R.
通讯作者: Stratton, Michael R.
DOI: 10.1038/nm.3984
发表时间: 2016-01
期刊: Nature medicine
影响因子: 82.9
作者:
Andor N;Graham TA;Jansen M;Xia LC;Aktipis CA;Petritsch C;Ji HP;Maley CC
通讯作者: Maley CC
DOI: 10.1038/nmeth0411-311
发表时间: 2011-04-01
期刊: NATURE METHODS
影响因子: 48
作者:
Kalisky, Tomer;Quake, Stephen R.
通讯作者: Quake, Stephen R.
DOI: 10.18632/oncotarget.7451
发表时间: 2016-03-15
期刊: ONCOTARGET
影响因子: --
作者:
Budczies, Jan;Pfarr, Nicole;Denkert, Carsten
通讯作者: Denkert, Carsten