BinPacker: Packing-Based De Novo Transcriptome Assembly from RNA-seq Data.

BinPacker: Packing-Based De Novo Transcriptome Assembly from RNA-seq Data.
复制标题

BinPacker:基于 RNA-seq 数据的包装 De Novo 转录组组装

DOI:
10.1371/journal.pcbi.1004772
复制
发表时间:
2016-02
影响因子:
4.3
通讯作者:
Huang X
Huang X
中科院分区:
生物学2区
文献类型:
--
作者:
Liu J;Li G;Chang Z;Yu T;Liu B;McMullen R;Chen P;Huang X

文献摘要

被引文献

相似文献

高通量RNA-seq技术为揭示转录组非常复杂的结构提供了前所未有的机会。然而,将大量的短RNA-seq读段组装成具有选择性剪接异构体的转录组是一项重要且极具挑战性的任务。在这项研究中,我们提出了一种新的从头组装,BinPacker,通过建模的转录组组装问题,跟踪一组轨迹的项目,其大小代表其相应的异构体的覆盖范围,通过解决一系列装箱问题。这种方法巧妙地将覆盖信息集成到过程中,具有两个独特的特征:1)仅剪接结点参与组装过程; 2)通过沿着剪接图上的结点边缘移动梳状结构来组装大量混乱读段。在真实的和模拟的RNA-seq数据集上进行测试,它在所有测试数据集上的性能几乎超过了所有现有的从头组装器,甚至在真实的狗数据集上的性能超过了从头组装器。此外,它比大多数汇编器运行得更快,需要的内存空间更少。BinPacker是在GNU通用公共许可证下发布的,其源代码可从http://sourceforge.net/projects/transcriptomeassembly/files/BinPacker_1.0.tar.gz/download获得。快速安装版本可从http://sourceforge.net/projects/transcriptomeassembly/files/BinPacker_binary.tar.gz/download获得。RNA-seq技术的可用性推动了从非常短的RNA序列进行转录组组装的算法的开发。然而,如何使用RNA-seq数据集(从头)组装转录组的问题尚未得到很好的建模;例如,序列覆盖信息甚至没有准确有效地整合到适当的组装程序中,导致所有现有(从头)策略都遇到了瓶颈。我们提出了一种新的方法来改造的问题,跟踪一组轨迹的项目,其大小代表其相应的异构体的覆盖范围,通过解决一系列装箱问题。这种方法巧妙地将覆盖信息集成到过程中,具有两个独特的特征:1)仅拼接结点参与组装过程; 2)通过沿着拼接图上的结点边缘移动梳状结构来组装大量混乱读段。在真实的和模拟的RNA-seq数据集上进行测试,就常用的比较标准而言,它在所有测试数据集上的性能都优于几乎所有现有的从头组装器,甚至优于那些从头组装器。
High-throughput RNA-seq technology has provided an unprecedented opportunity to reveal the very complex structures of transcriptomes. However, it is an important and highly challenging task to assemble vast amounts of short RNA-seq reads into transcriptomes with alternative splicing isoforms. In this study, we present a novel de novo assembler, BinPacker, by modeling the transcriptome assembly problem as tracking a set of trajectories of items with their sizes representing coverage of their corresponding isoforms by solving a series of bin-packing problems. This approach, which subtly integrates coverage information into the procedure, has two exclusive features: 1) only splicing junctions are involved in the assembling procedure; 2) massive pell-mell reads are assembled seemingly by moving a comb along junction edges on a splicing graph. Being tested on both real and simulated RNA-seq datasets, it outperforms almost all the existing de novo assemblers on all the tested datasets, and even outperforms those ab initio assemblers on the real dog dataset. In addition, it runs substantially faster and requires less memory space than most of the assemblers. BinPacker is published under GNU GENERAL PUBLIC LICENSE and the source is available from: http://sourceforge.net/projects/transcriptomeassembly/files/BinPacker_1.0.tar.gz/download. Quick installation version is available from: http://sourceforge.net/projects/transcriptomeassembly/files/BinPacker_binary.tar.gz/download. The availability of RNA-seq technology drives the development of algorithms for transcriptome assembly from very short RNA sequences. However, the problem of how to (de novo) assemble transcriptome using RNA-seq datasets has not been modeled well; e.g. sequence coverage information has even not been accurately and effectively integrated into the appropriate assembling procedure, leading to a bottleneck that all the existing (de novo) strategies have encountered. We present a novel approach to remodel the problem as tracking a set of trajectories of items with their sizes representing the coverage of their corresponding isoforms by solving a series of bin-packing problems. This approach, which subtly integrates the coverage information into the procedure, has two exclusive features: 1) only splicing junctions are involved in the assembling procedure; 2) massive pell-mell reads are assembled seemingly by moving a comb along junction edges on a splicing graph. Being tested on both real and simulated RNA-seq datasets, it outperforms almost all existing de novo assemblers on all the tested datasets, even outperforms those ab initio assemblers on the dog dataset, in terms of commonly used comparison standards.