Using the "Hidden" genome to improve classification of cancer types.

Using the "Hidden" genome to improve classification of cancer types.
复制标题

使用“隐藏”基因组改善癌症类型的分类。

DOI:
10.1111/biom.13367
复制
发表时间:
2021-12
期刊:
影响因子:
1.9
通讯作者:
Shen R
Shen R
中科院分区:
数学3区
文献类型:
--
作者:
Chakraborty S;Begg CB;Shen R

文献摘要

被引文献

相似文献

在临床上,使用识别体细胞突变的技术来检查癌症标本越来越常见。原则上,这些突变图谱可以用来诊断起源组织,对于原发部位未知的3%-5%的肿瘤来说,这是一项关键任务。原发部位的诊断对于使用循环DNA的筛查试验也是至关重要的。然而,在任何新的肿瘤中观察到的大多数突变都是非常罕见的突变,事实上,这些突变的优势可能从未在任何以前记录的肿瘤中观察到。为了创造一个可行的诊断工具,我们需要利用这个“隐藏的基因组”中的信息内容,这些变异没有直接的信息可用。为了实现这一点,我们提出了一种多层元特征回归方法来从训练数据中的稀有变量中提取关键信息,这种方法允许我们从新的肿瘤样本中任何以前未观察到的变量中提取诊断信息。通过将高维特征筛选方法与基于多层模型的等价混合效应表示的组套索惩罚最大似然方法相结合,获得了该模型的可扩展实现。我们将该方法应用于癌症基因组图谱全外显子组测序数据集,该数据集包括7个常见癌症部位的3702个肿瘤样本。结果表明,我们的多层次方法可以利用隐藏的基因组中的大量诊断信息。
It is increasingly common clinically for cancer specimens to be examined using techniques that identify somatic mutations. In principle these mutational profiles can be used to diagnose the tissue of origin, a critical task for the 3–5% of tumors that have an unknown primary site. Diagnosis of primary site is also critical for screening tests that employ circulating DNA. However, most mutations observed in any new tumor are very rarely occurring mutations, and indeed the preponderance of these may never have been observed in any previous recorded tumor. To create a viable diagnostic tool we need to harness the information content in this “hidden genome” of variants for which no direct information is available. To accomplish this we propose a multi-level meta-feature regression to extract the critical information from rare variants in the training data in a way that permits us to also extract diagnostic information from any previously unobserved variants in the new tumor sample. A scalable implementation of the model is obtained by combining a high-dimensional feature screening approach with a group-lasso penalized maximum likelihood approach based on an equivalent mixed-effect representation of the multilevel model. We apply the method to the Cancer Genome Atlas whole-exome sequencing data set including 3702 tumor samples across 7 common cancer sites. Results show that our multi-level approach can harness substantial diagnostic information from the hidden genome.