Investigating open reading frames in known and novel transcripts using ORFanage

Investigating open reading frames in known and novel transcripts using ORFanage
复制标题

使用 ORFanage 研究已知和新颖转录本中的开放阅读框

DOI:
10.1038/s43588-023-00496-1
复制
发表时间:
2023
期刊:
Nature Computational Science
影响因子:
--
通讯作者:
Pertea, Mihaela
Pertea, Mihaela
中科院分区:
--
文献类型:
--
作者:
Varabyou, Ales;Erdogdu, Beril;Salzberg, Steven L.;Pertea, Mihaela

文献摘要

相似文献

ORFanage是一个系统,旨在分配开放阅读框架(ORF)的已知和新的基因转录本,同时最大限度地提高相似性注释的蛋白质。ORFanage的主要用途是在RNA测序实验的组装结果中鉴定ORF,这是大多数转录组组装方法所不具备的能力。我们的实验展示了ORFanage如何用于在RNA-seq数据集中发现新的蛋白质变体,以及如何在人类注释数据库中的数万个转录本模型中改进ORF的注释。通过实现高度准确和高效的伪比对算法,ORFanage比其他ORF注释方法快得多,使其能够应用于非常大的数据集。当用于分析转录组组装时,ORFanage可以帮助从转录噪音中分离信号并识别可能的功能性转录变体,最终促进我们对生物学和医学的理解。
ORFanage is a system designed to assign open reading frames (ORFs) to known and novel gene transcripts while maximizing similarity to annotated proteins. The primary intended use of ORFanage is the identification of ORFs in the assembled results of RNA-sequencing experiments, a capability that most transcriptome assembly methods do not have. Our experiments demonstrate how ORFanage can be used to find novel protein variants in RNA-seq datasets, and to improve the annotations of ORFs in tens of thousands of transcript models in the human annotation databases. Through its implementation of a highly accurate and efficient pseudo-alignment algorithm, ORFanage is substantially faster than other ORF annotation methods, enabling its application to very large datasets. When used to analyze transcriptome assemblies, ORFanage can aid in the separation of signal from transcriptional noise and the identification of likely functional transcript variants, ultimately advancing our understanding of biology and medicine.