APA-Scan: detection and visualization of 3'-UTR alternative polyadenylation with RNA-seq and 3'-end-seq data.

APA-Scan: detection and visualization of 3'-UTR alternative polyadenylation with RNA-seq and 3'-end-seq data.
复制标题

DOI:
10.1186/s12859-022-04939-w
复制
发表时间:
2022-09-28
期刊:
影响因子:
3
通讯作者:
Zhang, Wei
Zhang, Wei
中科院分区:
生物学4区
文献类型:
--
作者:
Fahmi, Naima Ahmed;Ahmed, Khandakar Tanvir;Chang, Jae-Woong;Nassereddeen, Heba;Fan, Deliang;Yong, Jeongsik;Zhang, Wei

文献摘要

参考文献

被引文献

相似文献

真核基因组能够在前 mRNA 加工过程中通过选择性多聚腺苷酸化 (APA) 从基因产生多种同种型。 mRNA 3'-非翻译区 (3'-UTR) 中的 APA 产生具有更短或更长 3'-UTR 的转录物。通常,3'-UTR 充当 microRNA 和 RNA 结合蛋白的结合平台,影响 mRNA 转录物的命运。因此,已知 3'-UTR APA 可以调节翻译并提供在转录后水平调节基因表达的方法。由于注释不完整和分析能力低,当前的生物信息学流程在分析 3'-UTR APA 事件方面能力有限:广泛使用的生物信息学流程不引用可操作的聚腺苷酸化(切割)位点,而是仅使用 RNA-seq 读取覆盖来模拟 3'-UTR APA,从而导致误报识别。为了克服这些限制,我们开发了 APA-Scan,这是一个强大的程序,可以识别 3'-UTR APA 事件,并通过基因注释可视化 RNA-seq 短读覆盖范围。 APA-Scan 利用预测或实验验证的可操作聚腺苷酸化信号作为聚腺苷酸化位点的参考,并计算 RNA-seq 数据中长和短 3'-UTR 转录本的数量。 APA-Scan 的工作原理分为三个主要步骤:(i) 计算基因 3'-UTR 区域的读取覆盖率; (ii) 确定潜在的 APA 位点并评估两种生物条件下事件的重要性; (iii) 用户特定事件的图形表示,具有 3'-UTR 注释和 3'-UTR 区域的读数覆盖率。 APA-Scan 在 Python3 中实现。源代码和综合用户手册可在 https://github.com/compbiolabucf/APA-Scan 上免费获取。 APA-Scan 应用于模拟和真实 RNA-seq 数据集,并与两个广泛使用的基线 DaPars 和 APAtrap 进行比较。在模拟中,与其他基线相比,APA-Scan 显着提高了 3'-UTR APA 识别的准确性。 APA-Scan 的性能还通过小鼠胚胎成纤维细胞的 3'-end-seq 数据和 qPCR 进行了验证。实验证实 APA-Scan 可以检测未注释的 3'-UTR APA 事件并改善基因组注释。 APA-Scan 是一个综合计算管道,用于检测转录组范围内的 3'-UTR APA 事件。该流程集成了 RNA-seq 和 3'-end-seq 数据信息,可以通过高分辨率短读长覆盖图有效识别重要事件。在线版本包含可在 10.1186/s12859-022-04939-w 获取的补充材料。
The eukaryotic genome is capable of producing multiple isoforms from a gene by alternative polyadenylation (APA) during pre-mRNA processing. APA in the 3′-untranslated region (3′-UTR) of mRNA produces transcripts with shorter or longer 3′-UTR. Often, 3′-UTR serves as a binding platform for microRNAs and RNA-binding proteins, which affect the fate of the mRNA transcript. Thus, 3′-UTR APA is known to modulate translation and provides a mean to regulate gene expression at the post-transcriptional level. Current bioinformatics pipelines have limited capability in profiling 3′-UTR APA events due to incomplete annotations and a low-resolution analyzing power: widely available bioinformatics pipelines do not reference actionable polyadenylation (cleavage) sites but simulate 3′-UTR APA only using RNA-seq read coverage, causing false positive identifications. To overcome these limitations, we developed APA-Scan, a robust program that identifies 3′-UTR APA events and visualizes the RNA-seq short-read coverage with gene annotations. APA-Scan utilizes either predicted or experimentally validated actionable polyadenylation signals as a reference for polyadenylation sites and calculates the quantity of long and short 3′-UTR transcripts in the RNA-seq data. APA-Scan works in three major steps: (i) calculate the read coverage of the 3′-UTR regions of genes; (ii) identify the potential APA sites and evaluate the significance of the events among two biological conditions; (iii) graphical representation of user specific event with 3′-UTR annotation and read coverage on the 3′-UTR regions. APA-Scan is implemented in Python3. Source code and a comprehensive user’s manual are freely available at https://github.com/compbiolabucf/APA-Scan. APA-Scan was applied to both simulated and real RNA-seq datasets and compared with two widely used baselines DaPars and APAtrap. In simulation APA-Scan significantly improved the accuracy of 3′-UTR APA identification compared to the other baselines. The performance of APA-Scan was also validated by 3′-end-seq data and qPCR on mouse embryonic fibroblast cells. The experiments confirm that APA-Scan can detect unannotated 3′-UTR APA events and improve genome annotation. APA-Scan is a comprehensive computational pipeline to detect transcriptome-wide 3′-UTR APA events. The pipeline integrates both RNA-seq and 3′-end-seq data information and can efficiently identify the significant events with a high-resolution short reads coverage plots. The online version contains supplementary material available at 10.1186/s12859-022-04939-w.
DOI: 10.1093/bioinformatics/btv035
发表时间: 2015-06-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Le Pera L;Mazzapioda M;Tramontano A
通讯作者: Tramontano A
替代性裂解和聚腺苷酸化:长而短。
DOI: 10.1016/j.tibs.2013.03.005
发表时间: 2013-06
影响因子: 13.8
作者:
Tian, Bin;Manley, James L.
通讯作者: Manley, James L.
DOI: 10.1371/journal.pgen.1005879
发表时间: 2016-02
期刊: PLoS genetics
影响因子: 4.5
作者:
Hoffman Y;Bublik DR;Ugalde AP;Elkon R;Biniashvili T;Agami R;Oren M;Pilpel Y
通讯作者: Pilpel Y
DOI: 10.1016/j.cell.2009.06.016
发表时间: 2009-08-21
期刊: Cell
影响因子: 64.5
作者:
Mayr C;Bartel DP
通讯作者: Bartel DP
DOI: 10.1261/rna.2581711
发表时间: 2011-04-01
期刊: RNA
影响因子: 4.5
作者:
Shepard, Peter J.;Choi, Eun-A;Shi, Yongsheng
通讯作者: Shi, Yongsheng