lncScore: alignment-free identification of long noncoding RNA from assembled novel transcripts.

lncScore: alignment-free identification of long noncoding RNA from assembled novel transcripts.
复制标题

lncScore:从组装的新转录本中对长非编码RNA进行免比对鉴定

DOI:
10.1038/srep34838
复制
发表时间:
2016-10-06
期刊:
影响因子:
4.6
通讯作者:
Wang K
Wang K
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Zhao J;Song X;Wang K

文献摘要

参考文献

被引文献

相似文献

基于RNA-Seq的转录组组装已被广泛用于识别新的lncRNAs。然而,表现最好的转录重建方法只识别了智人全长蛋白质编码转录本的21%。这些部分长度的蛋白质编码转录本由于其CDS不完整而更有可能被归类为lncRNA,导致对lncRNA识别的假阳性率较高。此外,获得或取消终止密码子的潜在测序或组装错误也使基于ORF的lncRNAs预测复杂化。因此,从组装的转录本中识别lncRNA仍然是一个挑战,特别是部分长度的转录本。在这里,我们提出了一个新的无对齐工具,LncScore,它使用了11个精心挑选的特征的Logistic回归模型。与其他最先进的非比对工具(如Cpat、CNCI和PLEK)相比,LncScore在准确区分lncRNAs和mRNAs方面优于它们,特别是在人类和小鼠数据集中的部分长度mRNAs。此外,LncScore在其他五个物种(斑马鱼、苍蝇、线虫、老鼠和绵羊)的转录上也表现良好。为了加快预测速度,在lncScore内实现了多线程,分类64,756份转录只需2 分钟,训练12个线程的21,000份转录的新模型只需54 秒,比其他工具快得多。LncScore可在https://github.com/WGLab/lncScore.上获得
RNA-Seq based transcriptome assembly has been widely used to identify novel lncRNAs. However, the best-performing transcript reconstruction methods merely identified 21% of full-length protein-coding transcripts from H. sapiens. Those partial-length protein-coding transcripts are more likely to be classified as lncRNAs due to their incomplete CDS, leading to higher false positive rate for lncRNA identification. Furthermore, potential sequencing or assembly error that gain or abolish stop codons also complicates ORF-based prediction of lncRNAs. Therefore, it remains a challenge to identify lncRNAs from the assembled transcripts, particularly the partial-length ones. Here, we present a novel alignment-free tool, lncScore, which uses a logistic regression model with 11 carefully selected features. Compared to other state-of-the-art alignment-free tools (e.g. CPAT, CNCI, and PLEK), lncScore outperforms them on accurately distinguishing lncRNAs from mRNAs, especially partial-length mRNAs in the human and mouse datasets. In addition, lncScore also performed well on transcripts from five other species (Zebrafish, Fly, C. elegans, Rat, and Sheep). To speed up the prediction, multithreading is implemented within lncScore, and it only took 2 minute to classify 64,756 transcripts and 54 seconds to train a new model with 21,000 transcripts with 12 threads, which is much faster than other tools. lncScore is available at https://github.com/WGLab/lncScore.
DOI: 10.4161/epi.20170
发表时间: 2012-06-01
期刊: Epigenetics
影响因子: 3.7
作者:
Blignaut M
通讯作者: Blignaut M
DOI: 10.1101/gr.135350.111
发表时间: 2012-09
期刊: Genome research
影响因子: 7
作者:
Harrow J;Frankish A;Gonzalez JM;Tapanari E;Diekhans M;Kokocinski F;Aken BL;Barrell D;Zadissa A;Searle S;Barnes I;Bignell A;Boychenko V;Hunt T;Kay M;Mukherjee G;Rajan J;Despacio-Reyes G;Saunders G;Steward C;Harte R;Lin M;Howald C;Tanzer A;Derrien T;Chrast J;Walters N;Balasubramanian S;Pei B;Tress M;Rodriguez JM;Ezkurdia I;van Baren J;Brent M;Haussler D;Kellis M;Valencia A;Reymond A;Gerstein M;Guigó R;Hubbard TJ
通讯作者: Hubbard TJ
DOI: 10.1093/bioinformatics/btr209
发表时间: 2011-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lin MF;Jungreis I;Kellis M
通讯作者: Kellis M
DOI: 10.1093/bioinformatics/btv480
发表时间: 2015-12-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Achawanantakun, Rujira;Chen, Jiao;Zhang, Yuan
通讯作者: Zhang, Yuan
DOI: 10.1261/rna.047324.114
发表时间: 2015-03
期刊: RNA (New York, N.Y.)
影响因子: --
作者:
Haerty W;Ponting CP
通讯作者: Ponting CP