biobambam: tools for read pair collation based algorithms on BAM files

biobambam: tools for read pair collation based algorithms on BAM files
复制标题

biobambam:用于 BAM 文件上基于读取对校对的算法的工具

DOI:
10.1186/1751-0473-9-13
复制
发表时间:
2014-06-20
影响因子:
--
通讯作者:
Leonard S
Leonard S
中科院分区:
其他
文献类型:
--
作者:
Tischler G;Leonard S

文献摘要

被引文献

相似文献

当存储在BAM文件中时,序列比对数据通常按坐标(参考序列的id加上片段被映射的序列上的位置)排序,因为这简化了映射数据和参考之间的变体或映射数据内的变体的提取。按照这种顺序,配对读段通常在文件中分开,这使得一些其他应用复杂化,例如重复标记或转换为FastQ格式,这些应用需要访问配对的全部信息。在本文中,我们介绍biobambam,一套工具的基础上,有效的整理BAM文件中的路线读名称。所采用的排序算法避免了通过读段名称对比对进行耗时和空间消耗的排序,其中这在不使用超过指定量的主存储器的情况下是可能的。使用此算法,可以在有限的资源下非常有效地执行BAM文件中的重复标记和BAM文件到FastQ格式的转换等任务。我们还将排序规则算法以API的形式提供给其他项目。这个API是libmaus包的一部分。与以前的方法相比,涉及通过读名称(如BAM到FastQ或重复标记实用程序)整理比对的问题,我们的方法通常可以在所需的主内存和运行时间方面更有效地执行等效任务。我们的BAM到FastQ的转换速度比包括Picard和bamUtil在内的所有众所周知的替代方案都要快。对于小数据集,我们的重复标记与最接近的竞争对手bamUtil一样快,并且在大型和复杂数据集上比所有已知的替代方案更快。
Sequence alignment data is often ordered by coordinate (id of the reference sequence plus position on the sequence where the fragment was mapped) when stored in BAM files, as this simplifies the extraction of variants between the mapped data and the reference or of variants within the mapped data. In this order paired reads are usually separated in the file, which complicates some other applications like duplicate marking or conversion to the FastQ format which require to access the full information of the pairs. In this paper we introduce biobambam, a set of tools based on the efficient collation of alignments in BAM files by read name. The employed collation algorithm avoids time and space consuming sorting of alignments by read name where this is possible without using more than a specified amount of main memory. Using this algorithm tasks like duplicate marking in BAM files and conversion of BAM files to the FastQ format can be performed very efficiently with limited resources. We also make the collation algorithm available in the form of an API for other projects. This API is part of the libmaus package. In comparison with previous approaches to problems involving the collation of alignments by read name like the BAM to FastQ or duplication marking utilities our approach can often perform an equivalent task more efficiently in terms of the required main memory and run-time. Our BAM to FastQ conversion is faster than all widely known alternatives including Picard and bamUtil. Our duplicate marking is about as fast as the closest competitor bamUtil for small data sets and faster than all known alternatives on large and complex data sets.