Contrasting and Combining Transcriptome Complexity Captured by Short and Long RNA Sequencing Reads.

Contrasting and Combining Transcriptome Complexity Captured by Short and Long RNA Sequencing Reads.
复制标题

对比和组合短和长 RNA 测序读取捕获的转录组复杂性。

DOI:
10.1101/2023.11.21.568046
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Barash,Yoseph
Barash,Yoseph
中科院分区:
--
文献类型:
--
作者:
Han,SeongWoo;Jewell,San;Thomas-Tikhonenko,Andrei;Barash,Yoseph

文献摘要

相似文献

使用短读或长读RNA测序绘制转录变异图是基因组研究的主要内容。长读数能够捕获整个亚型并克服重复区域,而短读数仍然提供更好的覆盖率和错误率。然而,仍然存在悬而未决的问题,例如如何量化比较这些技术,我们能否将它们结合起来,以及这种结合的观点有什么好处?我们首先创建一个渠道,使用各种转录组统计数据来评估匹配的长读和短读数据,从而解决这些问题。我们发现,在数据集、算法和技术中,匹配的短读数据检测到的剪接连接比∼多30%,因此∼20%短读包含的剪接连接中有10%-30%被长读数据遗漏。相比之下,长阅读可以检测到更多的内含子保留事件,并可以检测到完整的异构体,这表明了结合这些技术的好处。我们引入了MAJIQ-L,它是MAJIQ软件的一个扩展,使人们能够统一地查看两种技术的转录组变异,并展示了它的好处。我们的软件可以用来评估任何未来的长阅读技术或算法,并可以与短阅读数据相结合,以改进转录组分析。
Mapping transcriptomic variations using either short- or long-read RNA sequencing is a staple of genomic research. Long reads are able to capture entire isoforms and overcome repetitive regions, whereas short reads still provide improved coverage and error rates. Yet, open questions remain, such as how to quantitatively compare the technologies, can we combine them, and what is the benefit of such a combined view? We tackle these questions by first creating a pipeline to assess matched long- and short-read data using a variety of transcriptome statistics. We find that across data sets, algorithms, and technologies, matched short-read data detects ∼30% more splice junctions, such that ∼10%–30% of the splice junctions included at ≥20% by short reads are missed by long reads. In contrast, long reads detect many more intron-retention events and can detect full isoforms, pointing to the benefit of combining the technologies. We introduce MAJIQ-L, an extension of the MAJIQ software, to enable a unified view of transcriptome variations from both technologies and demonstrate its benefits. Our software can be used to assess any future long-read technology or algorithm and can be combined with short-read data for improved transcriptome analysis.