LncRNA-ID: Long non-coding RNA IDentification using balanced random forests
LncRNA-ID: Long non-coding RNA IDentification using balanced random forests
复制标题
DOI:
10.1093/bioinformatics/btv480
复制
发表时间:
2015-12-15
期刊:
影响因子:
5.8
通讯作者:
Zhang, Yuan
中科院分区:
文献类型:
--
作者:
Achawanantakun, Rujira;Chen, Jiao;Zhang, Yuan
Motivation: Long non-coding RNAs (IncRNAs), which are non-coding RNAs of length above 200 nucleotides, play important biological functions such as gene expression regulation. To fully reveal the functions of IncRNAs, a fundamental step is to annotate them in various species. However, as IncRNAs tend to encode one or multiple open reading frames, it is not trivial to distinguish these long non-coding transcripts from protein-coding genes in transcriptomic data.Results: In this work, we design a new tool that calculates the coding potential of a transcript using a machine learning model (random forest) based on multiple features including sequence characteristics of putative open reading frames, translation scores based on ribosomal coverage, and conservation against characterized protein families. The experimental results show that our tool competes favorably with existing coding potential computation tools in IncRNA identification.Availability and implementation: The scripts and data can be downloaded at https://github.com/zhangy72/LncRNA-ID