Biological applications of support vector machines

Biological applications of support vector machines
复制标题

DOI:
10.1093/bib/5.4.328
复制
发表时间:
2004-12-01
影响因子:
9.5
通讯作者:
Yang, ZR
Yang, ZR
中科院分区:
生物学2区
文献类型:
--
作者:
Yang, ZR

文献摘要

被引文献

相似文献

生物信息学的主要任务之一是生物数据的分类和预测。随着生物数据库规模的快速增长,使用计算机程序来自动化分类过程是必不可少的。目前,给出最佳预测性能的计算机程序是支持向量机(SVM)。这是因为支持向量机的设计是为了最大化分隔两个类的裕度,以便训练的模型能够很好地概括看不见的数据。大多数其他计算机程序通过最小化训练中发生的错误来实现分类器,这导致泛化能力较差。正因为如此,支持向量机已被广泛应用于生物信息学的许多领域,包括蛋白质功能预测、蛋白酶功能位点识别、转录起始位点预测以及基因表达数据分类等。本文将讨论支持向量机的原理及其在生物数据分析中的应用,主要是蛋白质和DNA
One of the major tasks in bioinformatics is the classification and prediction of biological data. With the rapid increase in size of the biological databanks, it is essential to use computer programs to automate the classification process. At present, the computer programs that give the best prediction performance are support vector machines (SVMs). This is because SVMs are designed to maximise the margin to separate two classes so that the trained model generalises well on unseen data. Most other computer programs implement a classifier through the minimisation of error occurred in training, which leads to poorer generalisation. Because of this, SVMs have been widely applied to many areas of bioinformatics including protein function prediction, protease functional site recognition, transcription initiation site prediction and gene expression data classification. This paper will discuss the principles of SVMs and the applications of SVMs to the analysis of biological data, mainly protein and DNA