Distinguishing protein-coding from non-coding RNAs through support vector machines.

Distinguishing protein-coding from non-coding RNAs through support vector machines.
复制标题

DOI:
10.1371/journal.pgen.0020029
复制
发表时间:
2006-04
期刊:
影响因子:
4.5
通讯作者:
Rost B
Rost B
中科院分区:
生物学2区
文献类型:
--
作者:
Liu J;Gough J;Rost B

文献摘要

参考文献

被引文献

相似文献

RIKEN的FANTOM项目已经揭示了许多以前未知的编码序列,以及由替代性启动子使用和剪接导致的转录本的意想不到的变化程度。一般来说,转录组研究已经确定了更多不编码蛋白质的转录本。越来越多的证据表明这些非编码RNA(ncRNA)的重要细胞作用。因此,蛋白质编码RNA转录本与ncRNA转录本的区别是理解转录组和进行其注释的重要问题。很少有计算机方法专门解决这个问题。在这里,我们介绍了CONC(编码或非编码),一种基于支持向量机的新方法,该方法根据转录本编码蛋白质时所具有的特征对转录本进行分类。这些特征包括肽长度、氨基酸组成、预测的二级结构含量、预测的暴露残基百分比、组成熵、来自数据库搜索的同源物数量和比对熵。核苷酸频率也被并入该方法中。来自Swiss-Prot数据库的真核蛋白质的确认编码cDNA构成真阳性组,来自RNAdb和NONCODE的ncRNA构成真阴性组。10倍交叉验证表明,CONC区分编码RNA和ncRNA的特异性约为97%,灵敏度约为98%。应用于来自FANTOM 3数据集的102,801个小鼠cDNA,我们的方法可靠地鉴定了超过14,000个ncRNA,并估计ncRNA的总数约为28,000个。有两种类型的RNA:信使RNA(mRNA),它被翻译成蛋白质,和非编码RNA(ncRNA),它作为RNA分子的功能。除了教科书上的例子,如tRNA和rRNA,非编码RNA已被发现执行非常多样化的功能,从mRNA剪接和RNA修饰到翻译调控。据估计,非编码RNA占高等真核生物转录产物的绝大多数。从ncRNA中区分mRNA已经成为一个重要的生物学和计算问题。作者描述了一种基于机器学习算法的计算方法,称为支持向量机(SVM),该算法根据转录本编码蛋白质时所具有的特征对转录本进行分类。这些特征包括肽长度、氨基酸组成、二级结构含量和蛋白质比对信息。该方法应用于FANTOM 3大规模小鼠cDNA测序项目的数据集;它识别了小鼠中超过14,000个ncRNA,并估计FANTOM 3数据中的ncRNA总数约为28,000。
RIKEN's FANTOM project has revealed many previously unknown coding sequences, as well as an unexpected degree of variation in transcripts resulting from alternative promoter usage and splicing. Ever more transcripts that do not code for proteins have been identified by transcriptome studies, in general. Increasing evidence points to the important cellular roles of such non-coding RNAs (ncRNAs). The distinction of protein-coding RNA transcripts from ncRNA transcripts is therefore an important problem in understanding the transcriptome and carrying out its annotation. Very few in silico methods have specifically addressed this problem. Here, we introduce CONC (for “coding or non-coding”), a novel method based on support vector machines that classifies transcripts according to features they would have if they were coding for proteins. These features include peptide length, amino acid composition, predicted secondary structure content, predicted percentage of exposed residues, compositional entropy, number of homologs from database searches, and alignment entropy. Nucleotide frequencies are also incorporated into the method. Confirmed coding cDNAs for eukaryotic proteins from the Swiss-Prot database constituted the set of true positives, ncRNAs from RNAdb and NONCODE the true negatives. Ten-fold cross-validation suggested that CONC distinguished coding RNAs from ncRNAs at about 97% specificity and 98% sensitivity. Applied to 102,801 mouse cDNAs from the FANTOM3 dataset, our method reliably identified over 14,000 ncRNAs and estimated the total number of ncRNAs to be about 28,000. There are two types of RNA: messenger RNAs (mRNAs), which are translated into proteins, and non-coding RNAs (ncRNAs), which function as RNA molecules. Besides textbook examples such as tRNAs and rRNAs, non-coding RNAs have been found to carry out very diverse functions, from mRNA splicing and RNA modification to translational regulation. It has been estimated that non-coding RNAs make up the vast majority of transcription output of higher eukaryotes. Discriminating mRNA from ncRNA has become an important biological and computational problem. The authors describe a computational method based on a machine learning algorithm known as a support vector machine (SVM) that classifies transcripts according to features they would have if they were coding for proteins. These features include peptide length, amino acid composition, secondary structure content, and protein alignment information. The method is applied to the dataset from the FANTOM3 large-scale mouse cDNA sequencing project; it identifies over 14,000 ncRNAs in mouse and estimates the total number of ncRNAs in the FANTOM3 data to be about 28,000.
DOI: 10.1126/science.1112014
发表时间: 2005-09-02
期刊: SCIENCE
影响因子: 56.9
作者:
Carninci, P;Kasukawa, T;Hayashizaki, Y
通讯作者: Hayashizaki, Y
DOI: 10.1073/pnas.97.1.262
发表时间: 2000-01-04
影响因子: 11.1
作者:
Brown, MPS;Grundy, WN;Haussler, D
通讯作者: Haussler, D
DOI: 10.1093/nar/30.1.268
发表时间: 2002-01-01
影响因子: 14.9
作者:
Gough, J;Chothia, C
通讯作者: Chothia, C
DOI: 10.1093/embo-reports/kve230
发表时间: 2001-11-01
期刊: EMBO REPORTS
影响因子: 7.7
作者:
Mattick, JS
通讯作者: Mattick, JS
DOI: 10.1109/tnn.1997.641482
发表时间: 1997-01-01
影响因子: --
作者:
Cherkassky, V
通讯作者: Cherkassky, V