Massive computational identification of somatic variants in exonic splicing enhancers using The Cancer Genome Atlas

Massive computational identification of somatic variants in exonic splicing enhancers using The Cancer Genome Atlas
复制标题

使用癌症基因组图谱对外显子剪接增强子中的体细胞变异进行大规模计算鉴定

DOI:
10.1002/cam4.2619
复制
发表时间:
2019
期刊:
影响因子:
4
通讯作者:
Inazawa Johji
Inazawa Johji
中科院分区:
医学3区
文献类型:
--
作者:
Tanimoto Kousuke;Muramatsu Tomoki;Inazawa Johji

文献摘要

参考文献

相似文献

由于下一代测序(NGS)技术的发展,已经在各种类型的癌症中发现了大量的体细胞变异。然而,大多数体细胞变异的功能意义仍然未知。外显子剪接增强子(ESE)区域的体细胞变异被认为可以阻止富含丝氨酸和精氨酸(SR)的蛋白与ESE序列基序结合,从而导致外显子跳变。我们通过编译来自癌症基因组图谱(TCGA)的大量开放获取数据集,计算出ESEs的体细胞变异。利用32个TCGA项目中9635名患者的体细胞变异和RNA - seq数据,我们确定了646个ESE -干扰变异。我们的方法的假阳性率,估计使用排列测试,约为1%。在这些干扰ESE的变异中,大约71%位于四种经典SR蛋白的结合基序中。ESE‐干扰变异的发生与体细胞变异的数量成正比,但并不一定与癌症生物学过程相关的特定基因有关。现有的生物信息学工具无法预测本研究中发现的ESE‐干扰变异的致病性,尽管这些变异可能导致外显子跳变。我们证明了干扰ESE的无义变异倾向于逃避无义介导的衰变监测。通过对开放获取数据的综合分析,我们可以明确识别干扰ESE的变体。我们已经生成了一个强大的工具,它可以处理没有正常样本或原始数据的数据集,从而有助于减少不确定意义的变异,因为我们的统计方法只使用来自肿瘤样本的外显子连接读取计数。
Owing to the development of next‐generation sequencing (NGS) technologies, a large number of somatic variants have been identified in various types of cancer. However, the functional significance of most somatic variants remains unknown. Somatic variants that occur in exonic splicing enhancer (ESE) regions are thought to prevent serine and arginine‐rich (SR) proteins from binding to ESE sequence motifs, which leads to exon skipping. We computationally identified somatic variants in ESEs by compiling numerous open‐access datasets from The Cancer Genome Atlas (TCGA). Using somatic variants and RNA‐seq data from 9635 patients across 32 TCGA projects, we identified 646 ESE‐disrupting variants. The false positive rate of our method, estimated using a permutation test, was approximately 1%. Of these ESE‐disrupting variants, approximately 71% were located in the binding motifs of four classical SR proteins. ESE‐disrupting variants occurred in proportion to the number of somatic variants, but not necessarily in the specific genes associated with the biological processes of cancer. Existing bioinformatics tools could not predict the pathogenicity of ESE‐disrupting variants identified in this study, although these variants could cause exon skipping. We demonstrated that ESE‐disrupting nonsense variants tended to escape nonsense‐mediated decay surveillance. Using integrated analyses of open access data, we could specifically identify ESE‐disrupting variants. We have generated a powerful tool, which can handle datasets without normal samples or raw data, and thus contribute to reducing variants of uncertain significance because our statistical approach only uses the exon‐junction read counts from the tumor samples.
DOI: 10.1186/s13073-014-0121-3
发表时间: 2014
期刊: Genome medicine
影响因子: 12.3
作者:
Cheon JY;Mozersky J;Cook-Deegan R
通讯作者: Cook-Deegan R
DOI: 10.1093/nar/gkh393
发表时间: 2004-07-01
影响因子: 14.9
作者:
Fairbrother, WG;Yeo, GW;Burge, CB
通讯作者: Burge, CB
DOI: 10.1093/nar/gkw1108
发表时间: 2017-01-04
影响因子: 14.9
作者:
The Gene Ontology Consortium
通讯作者: The Gene Ontology Consortium
DOI: 10.1086/522036
发表时间: 2007-12-01
影响因子: 9.8
作者:
Conneely, Karen N.;Boehnke, Michael
通讯作者: Boehnke, Michael
DOI: 10.1016/j.ajhg.2016.08.016
发表时间: 2016-10-06
影响因子: 9.8
作者:
Ioannidis, Nilah M.;Rothstein, Joseph H.;Sieh, Weiva
通讯作者: Sieh, Weiva