Semi-supervised learning for classification of protein sequence data

Semi-supervised learning for classification of protein sequence data
复制标题

DOI:
10.3233/spr-2008-0241
复制
发表时间:
2008-01-01
影响因子:
--
通讯作者:
Guda, Chittibabu
Guda, Chittibabu
中科院分区:
计算机科学4区
文献类型:
--
作者:
King, Brian R.;Guda, Chittibabu

文献摘要

被引文献

相似文献

蛋白质序列数据继续以指数速度变得可用。这些数据的功能和结构属性的注释远远落后,只有一小部分数据被实验方法理解和标记。基于半监督学习的分类方法可以提高在许多领域中对部分标记数据进行分类的整体准确性,但很少有方法显示出它们对蛋白质序列分类的影响。我们展示了如何证明文本分类的方法可以应用于蛋白质序列数据,因为我们认为现有的和新的扩展的基本方法,并证明必须考虑的限制和差异。我们展示了比较结果对转导支持向量机,并显示最困难的分类问题上的上级结果。我们的研究结果表明,未标记蛋白质序列数据的大型存储库确实可以用于提高预测性能,特别是在标记蛋白质序列较少和/或数据本质上高度不平衡的情况下。
Protein sequence data continue to become available at an exponential rate. Annotation of functional and structural attributes of these data lags far behind, with only a small fraction of the data understood and labeled by experimental methods. Classification methods that are based on semi-supervised learning can increase the overall accuracy of classifying partly labeled data in many domains, but very few methods exist that have shown their effect on protein sequence classification. We show how proven methods from text classification can be applied to protein sequence data, as we consider both existing and novel extensions to the basic methods, and demonstrate restrictions and differences that must be considered. We demonstrate comparative results against the transductive support vector machine, and show superior results on the most difficult classification problems. Our results show that large repositories of unlabeled protein sequence data can indeed be used to improve predictive performance, particularly in situations where there are fewer labeled protein sequences available, and/or the data are highly unbalanced in nature.