L-RAPiT: A Cloud-Based Computing Pipeline for the Analysis of Long-Read RNA Sequencing Data.

L-RAPiT: A Cloud-Based Computing Pipeline for the Analysis of Long-Read RNA Sequencing Data.
复制标题

DOI:
10.3390/ijms232415851
复制
发表时间:
2022-12-13
影响因子:
5.6
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

已采用长读测序(LRS)来满足各种各样的研究需求,从构建新的转录组注释到快速鉴定新出现的病毒变体。在其他优点中,LRS在转录水平上比传统的高通量测序保留了更多关于RNA的信息,包括更准确和定量的剪接模式记录。使用LRS数据集的新研究正以指数速度发表,产生了大量的信息,可以用来解决许多不同的研究问题。然而,以量身定制的方式挖掘这种公开可用的数据目前并不容易,因为现有的软件工具通常需要熟悉命令行界面,这对许多研究人员构成了重大障碍。此外,不同的研究小组使用不同的软件包来执行LRS分析,这通常会阻止直接比较不同研究的已发表结果。为了应对这些挑战,我们开发了转录组学长读分析管道(L-RAPiT),这是一个用户友好的免费管道,不需要专门的计算资源或生物信息学专业知识。L-RAPiT可以直接通过Google Colaboratory实现,这是一个基于开源的Notebook环境的系统,可以直接分析来自Oxford Nanopore和PacBio LRS机器的转录组读数。这个新的管道可以快速,方便和标准化地分析公共可用或新生成的LRS数据集。
Long-read sequencing (LRS) has been adopted to meet a wide variety of research needs, ranging from the construction of novel transcriptome annotations to the rapid identification of emerging virus variants. Amongst other advantages, LRS preserves more information about RNA at the transcript level than conventional high-throughput sequencing, including far more accurate and quantitative records of splicing patterns. New studies with LRS datasets are being published at an exponential rate, generating a vast reservoir of information that can be leveraged to address a host of different research questions. However, mining such publicly available data in a tailored fashion is currently not easy, as the available software tools typically require familiarity with the command-line interface, which constitutes a significant obstacle to many researchers. Additionally, different research groups utilize different software packages to perform LRS analysis, which often prevents a direct comparison of published results across different studies. To address these challenges, we have developed the Long-Read Analysis Pipeline for Transcriptomics (L-RAPiT), a user-friendly, free pipeline requiring no dedicated computational resources or bioinformatics expertise. L-RAPiT can be implemented directly through Google Colaboratory, a system based on the open-source Jupyter notebook environment, and allows for the direct analysis of transcriptomic reads from Oxford Nanopore and PacBio LRS machines. This new pipeline enables the rapid, convenient, and standardized analysis of publicly available or newly generated LRS datasets.