Identification of gene fusions from human lung cancer mass spectrometry data.

Identification of gene fusions from human lung cancer mass spectrometry data.
复制标题

DOI:
10.1186/1471-2164-14-s8-s5
复制
发表时间:
2013
期刊:
影响因子:
4.4
通讯作者:
Xie L
Xie L
中科院分区:
生物学2区
文献类型:
--
作者:
Sun H;Xing X;Li J;Zhou F;Chen Y;He Y;Li W;Wei G;Chang X;Jia J;Li Y;Xie L

文献摘要

被引文献

相似文献

串联质谱学(MS/MS)技术已被应用于蛋白质鉴定,作为确认原始基因组注释的最终方法。为了能够识别基因融合蛋白,需要一个特殊的数据库,其中包含跨越基因融合断点的多肽。构建一个包含所有可能来源于潜在断裂点的融合肽的数据库是不切实际的。针对来自ChimerDB 2.0和癌症基因普查的6259个报告和预测的基因融合对,我们首次创建了一个数据库CanProFu,该数据库全面注释了这些配对基因之间的外显子-外显子连锁形成的融合肽。将该数据库应用于40例人非小细胞肺癌(NSCLC)和39例正常肺组织的质谱学数据集,在严格的搜索条件下,我们能够鉴定出19个独特的融合多肽,这些融合多肽表征了基因融合事件。其中11个基因融合事件仅在非小细胞肺癌组织中发现。此外,在癌变和正常肺标本中还发现了4种选择性剪接事件。这项工作中的数据库和工作流程可以灵活地应用于其他基于MS/MS的人类癌症实验,以检测作为潜在疾病生物标志物或药物靶点的基因融合。
Tandem mass spectrometry (MS/MS) technology has been applied to identify proteins, as an ultimate approach to confirm the original genome annotation. To be able to identify gene fusion proteins, a special database containing peptides that cross over gene fusion breakpoints is needed. It is impractical to construct a database that includes all possible fusion peptides originated from potential breakpoints. Focusing on 6259 reported and predicted gene fusion pairs from ChimerDB 2.0 and Cancer Gene Census, we for the first time created a database CanProFu that comprehensively annotates fusion peptides formed by exon-exon linkage between these pairing genes. Applying this database to mass spectrometry datasets of 40 human non-small cell lung cancer (NSCLC) samples and 39 normal lung samples with stringent searching criteria, we were able to identify 19 unique fusion peptides characterizing gene fusion events. Among them 11 gene fusion events were only found in NSCLC samples. And also, 4 alternative splicing events were characterized in cancerous or normal lung samples. The database and workflow in this work can be flexibly applied to other MS/MS based human cancer experiments to detect gene fusions as potential disease biomarkers or drug targets.