Comparative assessment of methods for the fusion transcripts detection from RNA-Seq data.

Comparative assessment of methods for the fusion transcripts detection from RNA-Seq data.
复制标题

DOI:
10.1038/srep21597
复制
发表时间:
2016-02-10
期刊:
影响因子:
4.6
通讯作者:
Li H
Li H
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Kumar S;Vo AD;Qin F;Li H

文献摘要

被引文献

相似文献

RNA-Seq 使融合转录本(即“嵌合 RNA”)的全局鉴定成为可能。尽管已经开发了各种软件包来实现此目的,但它们在不同开发人员提供的不同数据集中表现不同。对于用户和开发人员来说,对现有融合检测工具的性能进行公正的评估非常重要。为了这个目标,我们比较了 12 个知名融合检测软件包的性能。我们在四个不同的数据集(正数据集、负数据集、混合数据集和测试数据集)中评估了这些工具的灵敏度、错误发现率、计算时间和内存使用情况。我们得出的结论是,某些工具在灵敏度、正预测值、时间消耗和内存使用方面优于其他工具。我们还观察到真实数据集(测试数据集)中不同工具检测到的融合存在小的重叠。这可能是由于各种工具的错误发现,但也可能是因为没有一个工具具有包容性。我们发现工具的性能取决于 RNA-Seq 数据的质量、读段长度和读段数量。我们建议用户根据 RNA-Seq 数据的特性选择适合其目的的工具。
RNA-Seq made possible the global identification of fusion transcripts, i.e. “chimeric RNAs”. Even though various software packages have been developed to serve this purpose, they behave differently in different datasets provided by different developers. It is important for both users, and developers to have an unbiased assessment of the performance of existing fusion detection tools. Toward this goal, we compared the performance of 12 well-known fusion detection software packages. We evaluated the sensitivity, false discovery rate, computing time, and memory usage of these tools in four different datasets (positive, negative, mixed, and test). We conclude that some tools are better than others in terms of sensitivity, positive prediction value, time consumption and memory usage. We also observed small overlaps of the fusions detected by different tools in the real dataset (test dataset). This could be due to false discoveries by various tools, but could also be due to the reason that none of the tools are inclusive. We have found that the performance of the tools depends on the quality, read length, and number of reads of the RNA-Seq data. We recommend that users choose the proper tools for their purpose based on the properties of their RNA-Seq data.