Examining Tumor Phylogeny Inference in Noisy Sequencing Data

Examining Tumor Phylogeny Inference in Noisy Sequencing Data
复制标题

DOI:
10.1109/bibm.2018.8621437
复制
发表时间:
2018-12
期刊:
2018 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)
影响因子:
--
通讯作者:
K. Tomlinson;Layla Oesper
K. Tomlinson;Layla Oesper
中科院分区:
其他
文献类型:
--
作者:
K. Tomlinson;Layla Oesper

文献摘要

相似文献

最近已经提出了许多方法来从噪声DNA测序数据中重建肿瘤的进化历史。我们调查了当只考虑单核苷酸变体(SNV)时,何时以及如何从多样本批量测序数据中重建这些历史。我们将其形式化为枚举变量等位基因频率因式分解问题,并为与给定数据集一致的可能系统发育的数量的上限提供了一个新的证明。此外,我们提出并评估了两种方法来提高现有的基于图的系统发育推断方法的健壮性和性能。我们将我们的方法应用于有噪声的模拟数据,发现低覆盖率和高噪声使识别系统发育更加困难。我们还将我们的方法应用于慢性淋巴细胞性白血病和肾透明细胞癌数据集。
A number of methods have recently been proposed to reconstruct the evolutionary history of a tumor from noisy DNA sequencing data. We investigate when and how well these histories can be reconstructed from multi-sample bulk sequencing data when considering only single nucleotide variants (SNVs). We formalize this as the Enumeration Variant Allele Frequency Factorization Problem and provide a novel proof for an upper bound on the number of possible phylogenies consistent with a given dataset. In addition, we propose and assess two methods for increasing the robustness and performance of an existing graph based phylogenetic inference method. We apply our approaches to noisy simulated data and find that low coverage and high noise make it more difficult to identify phylogenies. We also apply our methods to both chronic lymphocytic leukemia and clear cell renal cell carcinoma datasets.