PSoL: a positive sample only learning algorithm for finding non-coding RNA genes

PSoL: a positive sample only learning algorithm for finding non-coding RNA genes
复制标题

DOI:
10.1093/bioinformatics/btl441
复制
发表时间:
2006-11-01
期刊:
影响因子:
5.8
通讯作者:
Holbrook, Stephen R.
Holbrook, Stephen R.
中科院分区:
生物学3区
文献类型:
--
作者:
Wang, Chunlin;Ding, Chris;Holbrook, Stephen R.

文献摘要

被引文献

相似文献

小分子非编码RNA(ncRNA)基因在多种细胞过程中起着重要的调控作用。然而,ncRNA基因的检测是一个巨大的挑战,无论是实验和计算方法。在这项研究中,我们描述了一种新的方法,称为阳性样本仅学习(PSoL)来预测大肠杆菌基因组中的ncRNA基因。虽然PSoL是一种用于分类的机器学习方法,但它不需要负训练数据,这通常很难正确定义,并且会显着影响机器学习的性能。此外,PSoL以支持向量机(SVM)为核心学习算法,可以融合多种不同类型的信息,提高预测精度。PSoL除了用于预测ncRNA外,还可以应用于其他生物信息学问题。结果:5倍交叉验证实验表明,PSoL在已知ncRNA的回收率上可以达到80%左右的准确率。我们将PSoL预测与之前发表的五个结果进行了比较。PSoL方法的预测与其他方法的预测重叠的百分比最高。
Motivation: Small non-coding RNA (ncRNA) genes play important regulatory roles in a variety of cellular processes. However, detection of ncRNA genes is a great challenge to both experimental and computational approaches. In this study, we describe a new approach called positive sample only learning (PSoL) to predict ncRNA genes in the Escherichia coli genome. Although PSoL is a machine learning method for classification, it requires no negative training data, which, in general, is hard to define properly and affects the performance of machine learning dramatically. In addition, using the support vector machine (SVM) as the core learning algorithm, PSoL can integrate many different kinds of information to improve the accuracy of prediction. Besides the application of PSoL for predicting ncRNAs, PSoL is applicable to many other bioinformatics problems as well.Results: The PSoL method is assessed by 5-fold cross-validation experiments which show that PSoL can achieve about 80% accuracy in recovery of known ncRNAs. We compared PSoL predictions with five previously published results. The PSoL method has the highest percentage of predictions overlapping with those from other methods.