LncRNA-ID: Long non-coding RNA IDentification using balanced random forests

LncRNA-ID: Long non-coding RNA IDentification using balanced random forests
复制标题

DOI:
10.1093/bioinformatics/btv480
复制
发表时间:
2015-12-15
期刊:
影响因子:
5.8
通讯作者:
Zhang, Yuan
Zhang, Yuan
中科院分区:
生物学3区
文献类型:
--
作者:
Achawanantakun, Rujira;Chen, Jiao;Zhang, Yuan

文献摘要

被引文献

相似文献

动机:长链非编码RNA(Long non-coding RNAs,IncRNA)是指长度超过200个核苷酸的非编码RNA,具有重要的生物学功能,如基因表达调控。为了充分揭示IncRNA的功能,一个基本的步骤是在不同物种中注释它们。然而,由于IncRNA倾向于编码一个或多个开放阅读框,因此在转录组数据中区分这些长的非编码转录本与蛋白质编码基因并非易事。在这项工作中,我们设计了一个新的工具,使用机器学习模型计算转录本的编码潜力(随机森林)基于多个特征,包括推定的开放阅读框的序列特征,基于核糖体覆盖度的翻译得分,和保守性。实验结果表明,我们的工具在IncRNA识别中与现有的编码势计算工具竞争有利。可用性和实现:脚本和数据可以在https://github.com/zhangy72/LncRNA-ID下载
Motivation: Long non-coding RNAs (IncRNAs), which are non-coding RNAs of length above 200 nucleotides, play important biological functions such as gene expression regulation. To fully reveal the functions of IncRNAs, a fundamental step is to annotate them in various species. However, as IncRNAs tend to encode one or multiple open reading frames, it is not trivial to distinguish these long non-coding transcripts from protein-coding genes in transcriptomic data.Results: In this work, we design a new tool that calculates the coding potential of a transcript using a machine learning model (random forest) based on multiple features including sequence characteristics of putative open reading frames, translation scores based on ribosomal coverage, and conservation against characterized protein families. The experimental results show that our tool competes favorably with existing coding potential computation tools in IncRNA identification.Availability and implementation: The scripts and data can be downloaded at https://github.com/zhangy72/LncRNA-ID