ChimPipe: accurate detection of fusion genes and transcription-induced chimeras from RNA-seq data.

ChimPipe: accurate detection of fusion genes and transcription-induced chimeras from RNA-seq data.
复制标题

DOI:
10.1186/s12864-016-3404-9
复制
发表时间:
2017-01-03
期刊:
影响因子:
4.4
通讯作者:
Djebali S
Djebali S
中科院分区:
生物学2区
文献类型:
--
作者:
Rodríguez-Martín B;Palumbo E;Marco-Sola S;Griebel T;Ribeca P;Alonso G;Rastrojo A;Aguado B;Guigó R;Djebali S

文献摘要

参考文献

被引文献

相似文献

嵌合转录物通常被定义为连接基因组中两个或更多个不同基因的转录物,并且可以通过各种生物学机制来解释,例如基因组重排、通读或反式剪接,但也可以通过技术或生物人工制品来解释。一些研究已经表明它们在癌症、细胞多能性和运动性中的重要性。最近开发了许多程序来从Illumina RNA-seq数据中识别嵌合体(主要是癌症中的融合基因)。然而,不同程序在同一数据集上的输出可能大不一致,并且往往包括许多误报。其他问题涉及限于融合基因的模拟数据集、具有有限数量的验证病例的真实的数据集、模拟数据集和真实的数据集之间的结果不一致以及基因而不是连接水平评估。在这里,我们提出了ChimPipe,这是一种模块化且易于使用的方法,可以从配对末端Illumina RNA-seq数据中可靠地识别融合基因和转录诱导的嵌合体。我们还为三种不同的读取长度产生了逼真的模拟数据集,并通过将确切的连接点与验证的基因融合相关联来增强两个金标准癌症数据集。ChimPipe与其他四种最先进的工具一起对这些数据进行基准测试,结果表明ChimPipe是识别两种数据集精确连接点坐标的顶级程序,并且是灵敏度和精度之间最佳权衡的程序。应用于106个ENCODE人类RNA-seq数据集,ChimPipe确定了137个连接其亲本基因的蛋白质编码序列的高置信度嵌合体。在随后的实验中,四分之三的预测嵌合体,其中两个在大多数样品中反复表达,可以验证。对这三个病例的克隆和测序揭示了几种新的嵌合转录物结构,其中3种具有编码嵌合蛋白的潜力,我们假设其具有新的作用。将ChimPipe应用于人类和小鼠的ENCODE RNA-seq数据,鉴定出131种两个物种共有的复发性嵌合体,因此可能是保守的。ChimPipe结合了不一致的配对末端读段和分裂读段来检测任何类型的嵌合体,包括那些源自聚合酶通读的嵌合体,并且在灵敏度和精确度之间显示出优异的权衡。ChimPipe发现的嵌合体可以在体外以高精度进行验证。本文的在线版本(doi:10.1186/s12864-016-3404-9)包含补充材料,可供授权用户使用。
Chimeric transcripts are commonly defined as transcripts linking two or more different genes in the genome, and can be explained by various biological mechanisms such as genomic rearrangement, read-through or trans-splicing, but also by technical or biological artefacts. Several studies have shown their importance in cancer, cell pluripotency and motility. Many programs have recently been developed to identify chimeras from Illumina RNA-seq data (mostly fusion genes in cancer). However outputs of different programs on the same dataset can be widely inconsistent, and tend to include many false positives. Other issues relate to simulated datasets restricted to fusion genes, real datasets with limited numbers of validated cases, result inconsistencies between simulated and real datasets, and gene rather than junction level assessment. Here we present ChimPipe, a modular and easy-to-use method to reliably identify fusion genes and transcription-induced chimeras from paired-end Illumina RNA-seq data. We have also produced realistic simulated datasets for three different read lengths, and enhanced two gold-standard cancer datasets by associating exact junction points to validated gene fusions. Benchmarking ChimPipe together with four other state-of-the-art tools on this data showed ChimPipe to be the top program at identifying exact junction coordinates for both kinds of datasets, and the one showing the best trade-off between sensitivity and precision. Applied to 106 ENCODE human RNA-seq datasets, ChimPipe identified 137 high confidence chimeras connecting the protein coding sequence of their parent genes. In subsequent experiments, three out of four predicted chimeras, two of which recurrently expressed in a large majority of the samples, could be validated. Cloning and sequencing of the three cases revealed several new chimeric transcript structures, 3 of which with the potential to encode a chimeric protein for which we hypothesized a new role. Applying ChimPipe to human and mouse ENCODE RNA-seq data led to the identification of 131 recurrent chimeras common to both species, and therefore potentially conserved. ChimPipe combines discordant paired-end reads and split-reads to detect any kind of chimeras, including those originating from polymerase read-through, and shows an excellent trade-off between sensitivity and precision. The chimeras found by ChimPipe can be validated in-vitro with high accuracy. The online version of this article (doi:10.1186/s12864-016-3404-9) contains supplementary material, which is available to authorized users.
DOI: 10.1093/nar/gkw032
发表时间: 2016-04-07
影响因子: 14.9
作者:
Babiceanu M;Qin F;Xie Z;Jia Y;Lopez K;Janus N;Facemire L;Kumar S;Pang Y;Qi Y;Lazar IM;Li H
通讯作者: Li H
DOI: 10.1101/gr.135350.111
发表时间: 2012-09
期刊: Genome research
影响因子: 7
作者:
Harrow J;Frankish A;Gonzalez JM;Tapanari E;Diekhans M;Kokocinski F;Aken BL;Barrell D;Zadissa A;Searle S;Barnes I;Bignell A;Boychenko V;Hunt T;Kay M;Mukherjee G;Rajan J;Despacio-Reyes G;Saunders G;Steward C;Harte R;Lin M;Howald C;Tanzer A;Derrien T;Chrast J;Walters N;Balasubramanian S;Pei B;Tress M;Rodriguez JM;Ezkurdia I;van Baren J;Brent M;Haussler D;Kellis M;Valencia A;Reymond A;Gerstein M;Guigó R;Hubbard TJ
通讯作者: Hubbard TJ
DOI: 10.1101/gr.4137606
发表时间: 2006-01-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Akiva, P;Toporik, A;Sorek, R
通讯作者: Sorek, R
DOI: 10.1186/1471-2164-14-199
发表时间: 2013-03-22
期刊: BMC GENOMICS
影响因子: 4.4
作者:
Hernandez-Torres, Francisco;Rastrojo, Alberto;Aguado, Begona
通讯作者: Aguado, Begona
DOI: 10.1101/gr.152132.112
发表时间: 2014-02
期刊: Genome research
影响因子: 7
作者:
Ferreira PG;Jares P;Rico D;Gómez-López G;Martínez-Trillos A;Villamor N;Ecker S;González-Pérez A;Knowles DG;Monlong J;Johnson R;Quesada V;Djebali S;Papasaikas P;López-Guerra M;Colomer D;Royo C;Cazorla M;Pinyol M;Clot G;Aymerich M;Rozman M;Kulis M;Tamborero D;Gouin A;Blanc J;Gut M;Gut I;Puente XS;Pisano DG;Martin-Subero JI;López-Bigas N;López-Guillermo A;Valencia A;López-Otín C;Campo E;Guigó R
通讯作者: Guigó R