A machine learning approach to optimizing cell-free DNA sequencing panels: with an application to prostate cancer

A machine learning approach to optimizing cell-free DNA sequencing panels: with an application to prostate cancer
复制标题

DOI:
10.1186/s12885-020-07318-x
复制
发表时间:
2020-08-28
期刊:
影响因子:
3.8
通讯作者:
Witte, John S.
Witte, John S.
中科院分区:
医学2区
文献类型:
--
作者:
Cario, Clinton L.;Chen, Emmalyn;Witte, John S.

文献摘要

被引文献

相似文献

由于恶性肿瘤的遗传异质性和肿瘤来源分子的稀有性,游离DNA (cfDNA)作为癌症生物标志物的应用具有挑战性。在这里,我们描述并展示了一种新的机器学习引导面板设计策略,用于改进cfDNA中肿瘤变异的检测。使用这种方法,我们首先生成了一个模型来对候选变异进行分类和评分,以便纳入前列腺癌靶向测序面板。然后,我们使用该小组在硅捕获和混合捕获设置中筛选局限性疾病前列腺癌患者的肿瘤变异。方法分析550例前列腺肿瘤的全基因组序列(WGS)数据,建立单点和小(< 200 bp) indel突变的靶向测序面板,随后与5例患者的前列腺肿瘤序列进行计算机筛选,评估其与常用替代面板设计的性能。该小组检测肿瘤来源的cfDNA变异的能力,然后使用前瞻性收集的cfDNA和来自18名接受根治性前列腺癌切除术的局限性前列腺癌患者的肿瘤灶进行评估。结果:该方法产生的小组确定了已知驱动基因(如HRAS)和前列腺癌相关转录因子结合位点(如MYC, AR)的突变为首选候选。当在计算机环境中分析5名前列腺癌患者的cfDNA时,它在检测体细胞突变方面优于两种常用的设计。此外,使用该面板对cfDNA分子进行混合捕获和2500X测序,在测试集的所有18例患者中检测到肿瘤变异,其中18例患者中有15例检测到在多个病灶中发现的变异。结论机器学习优先的靶向测序面板可能有助于广泛和敏感地检测异质性疾病的cfDNA变异。当应用于从前列腺癌患者分离的cfDNA时,该策略对疾病检测和监测具有意义。
Background Cell-free DNA's (cfDNA) use as a biomarker in cancer is challenging due to genetic heterogeneity of malignancies and rarity of tumor-derived molecules. Here we describe and demonstrate a novel machine-learning guided panel design strategy for improving the detection of tumor variants in cfDNA. Using this approach, we first generated a model to classify and score candidate variants for inclusion on a prostate cancer targeted sequencing panel. We then used this panel to screen tumor variants from prostate cancer patients with localized disease in both in silico and hybrid capture settings. Methods Whole Genome Sequence (WGS) data from 550 prostate tumors was analyzed to build a targeted sequencing panel of single point and small (< 200 bp) indel mutations, which was subsequently screened in silico against prostate tumor sequences from 5 patients to assess performance against commonly used alternative panel designs. The panel's ability to detect tumor-derived cfDNA variants was then assessed using prospectively collected cfDNA and tumor foci from a test set 18 prostate cancer patients with localized disease undergoing radical proctectomy. Results The panel generated from this approach identified as top candidates mutations in known driver genes (e.g. HRAS) and prostate cancer related transcription factor binding sites (e.g. MYC, AR). It outperformed two commonly used designs in detecting somatic mutations found in the cfDNA of 5 prostate cancer patients when analyzed in an in silico setting. Additionally, hybrid capture and 2500X sequencing of cfDNA molecules using the panel resulted in detection of tumor variants in all 18 patients of a test set, where 15 of the 18 patients had detected variants found in multiple foci. Conclusion Machine learning-prioritized targeted sequencing panels may prove useful for broad and sensitive variant detection in the cfDNA of heterogeneous diseases. This strategy has implications for disease detection and monitoring when applied to the cfDNA isolated from prostate cancer patients.