Predicting Cancer Risk from Germline Whole-exome Sequencing Data Using a Novel Context-based Variant Aggregation Approach.

Predicting Cancer Risk from Germline Whole-exome Sequencing Data Using a Novel Context-based Variant Aggregation Approach.
复制标题

DOI:
10.1158/2767-9764.crc-22-0355
复制
发表时间:
2023-03
期刊:
Cancer research communications
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

许多研究表明,肿瘤中体细胞变异的基因组、核苷酸和表观遗传背景的分布提供了癌症病因学的信息。最近,一个新的研究方向集中在从种系变异的背景中提取信号,并且已经出现了证据表明,由这些因素定义的模式与致癌途径,组织学亚型和预后相关。使用捕获其基因组、核苷酸和表观遗传背景的元特征聚合种系变异是否可以改善癌症风险预测仍然是一个悬而未决的问题。这种聚合方法可能会增加检测罕见变异信号的统计能力,这些变异被假设为癌症遗传性缺失的主要来源。使用来自英国生物库的生殖系全外显子组测序数据,我们使用已知的风险变体(已知癌症易感基因中的癌症相关SNP和致病性变体)以及另外包括元特征的模型开发了10种癌症类型的风险模型。元特征并没有提高基于已知风险变量的模型的预测准确性。将方法扩展到全基因组测序可能会导致预测准确性的提高。有证据表明,癌症部分是由尚未确定的罕见遗传变异引起的。我们使用新的统计方法和来自英国生物银行的数据来调查这个问题。
Many studies have shown that the distributions of the genomic, nucleotide, and epigenetic contexts of somatic variants in tumors are informative of cancer etiology. Recently, a new direction of research has focused on extracting signals from the contexts of germline variants and evidence has emerged that patterns defined by these factors are associated with oncogenic pathways, histologic subtypes, and prognosis. It remains an open question whether aggregating germline variants using meta-features capturing their genomic, nucleotide, and epigenetic contexts can improve cancer risk prediction. This aggregation approach can potentially increase statistical power for detecting signals from rare variants, which have been hypothesized to be a major source of the missing heritability of cancer. Using germline whole-exome sequencing data from the UK Biobank, we developed risk models for 10 cancer types using known risk variants (cancer-associated SNPs and pathogenic variants in known cancer predisposition genes) as well as models that additionally include the meta-features. The meta-features did not improve the prediction accuracy of models based on known risk variants. It is possible that expanding the approach to whole-genome sequencing can lead to gains in prediction accuracy. There is evidence that cancer is partly caused by rare genetic variants that have not yet been identified. We investigate this issue using novel statistical methods and data from the UK Biobank.