Prediction and classification of ncRNAs using structural information.

Prediction and classification of ncRNAs using structural information.
复制标题

DOI:
10.1186/1471-2164-15-127
复制
发表时间:
2014-02-13
期刊:
影响因子:
4.4
通讯作者:
Raghava GP
Raghava GP
中科院分区:
生物学2区
文献类型:
--
作者:
Panwar B;Arora A;Raghava GP

文献摘要

参考文献

被引文献

相似文献

越来越多的证据表明,非编码转录本,以前被认为是功能惰性,在各种细胞活动中发挥重要作用。像下一代测序这样的高通量技术已经导致了大量序列数据的产生。因此,不仅需要区分编码和非编码转录物,而且还需要将非编码RNA(ncRNA)转录物分配到相应的类别(家族)中。虽然有几种算法可用于此任务,其分类性能仍然是一个主要问题。认识到非编码转录本在细胞过程中发挥的关键作用,需要开发能够精确分类ncRNA转录本的算法。在这项研究中,我们首先开发预测工具来区分编码或非编码转录本,然后将ncRNA分为相应的类别。与采用多个特征的现有方法相比,我们基于SVM的方法通过使用单个特征(三核苷酸组成)实现了0.98的MCC。由于知道ncRNA转录本的结构可以提供对其生物学功能的见解,我们使用预测的ncRNA结构的图形属性将转录本分类为18个不同的非编码RNA类别。我们使用各种算法(BayeNet,NaiveBayes,MultilayerPerceptron,IBk,libSVM,SMO和RandomForest)开发了分类模型,并观察到基于RandomForest的模型比其他模型表现更好。与GraPPLE研究相比,灵敏度(13类)和特异性(14类)更高。此外,0.43的总体灵敏度优于GraPPLE的灵敏度(0.33),而0.40的总体MCC测量(与GraPPLE的0.29的MCC相比)对于我们的方法显著更高。这清楚地表明,我们的模型比现有模型更准确。这项工作最终证明,一个简单的功能,三核苷酸组成,是足以区分编码和非编码RNA序列。类似地,基于图属性的特征集沿着RandomForest算法最适合于分类不同的ncRNA类。我们还开发了一个在线独立工具- RNAcon(http://crdd.osdd.net/raghava/rnacon)。
Evidence is accumulating that non-coding transcripts, previously thought to be functionally inert, play important roles in various cellular activities. High throughput techniques like next generation sequencing have resulted in the generation of vast amounts of sequence data. It is therefore desirable, not only to discriminate coding and non-coding transcripts, but also to assign the noncoding RNA (ncRNA) transcripts into respective classes (families). Although there are several algorithms available for this task, their classification performance remains a major concern. Acknowledging the crucial role that non-coding transcripts play in cellular processes, it is required to develop algorithms that are able to precisely classify ncRNA transcripts. In this study, we initially develop prediction tools to discriminate coding or non-coding transcripts and thereafter classify ncRNAs into respective classes. In comparison to the existing methods that employed multiple features, our SVM-based method by using a single feature (tri-nucleotide composition), achieved MCC of 0.98. Knowing that the structure of a ncRNA transcript could provide insights into its biological function, we use graph properties of predicted ncRNA structures to classify the transcripts into 18 different non-coding RNA classes. We developed classification models using a variety of algorithms (BayeNet, NaiveBayes, MultilayerPerceptron, IBk, libSVM, SMO and RandomForest) and observed that model based on RandomForest performed better than other models. As compared to the GraPPLE study, the sensitivity (of 13 classes) and specificity (of 14 classes) was higher. Moreover, the overall sensitivity of 0.43 outperforms the sensitivity of GraPPLE (0.33) whereas the overall MCC measure of 0.40 (in contrast to MCC of 0.29 of GraPPLE) was significantly higher for our method. This clearly demonstrates that our models are more accurate than existing models. This work conclusively demonstrates that a simple feature, tri-nucleotide composition, is sufficient to discriminate between coding and non-coding RNA sequences. Similarly, graph properties based feature set along with RandomForest algorithm are most suitable to classify different ncRNA classes. We have also developed an online and standalone tool-- RNAcon ( http://crdd.osdd.net/raghava/rnacon).
DOI: 10.1126/science.1064921
发表时间: 2001-10-26
期刊: SCIENCE
影响因子: 56.9
作者:
Lagos-Quintana, M;Rauhut, R;Tuschl, T
通讯作者: Tuschl, T
DOI: 10.1093/nar/gkg006
发表时间: 2003-01-01
影响因子: 14.9
作者:
Griffiths-Jones, S;Bateman, A;Eddy, SR
通讯作者: Eddy, SR
DOI: 10.1016/j.sbi.2010.11.005
发表时间: 2011-02-01
影响因子: 6.8
作者:
Mason, Mark;Schuller, Anthony;Skordalakes, Emmanuel
通讯作者: Skordalakes, Emmanuel
DOI: 10.1371/journal.pgen.0020029
发表时间: 2006-04
期刊: PLoS genetics
影响因子: 4.5
作者:
Liu J;Gough J;Rost B
通讯作者: Rost B
DOI: 10.1038/nature07756
发表时间: 2009-01-22
期刊: NATURE
影响因子: 64.8
作者:
Moazed, Danesh
通讯作者: Moazed, Danesh