Knowledge-based analysis of microarray gene expression data by using support vector machines

Knowledge-based analysis of microarray gene expression data by using support vector machines
复制标题

DOI:
10.1073/pnas.97.1.262
复制
发表时间:
2000-01-04
影响因子:
11.1
通讯作者:
Haussler, D
Haussler, D
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Brown, MPS;Grundy, WN;Haussler, D

文献摘要

被引文献

相似文献

我们通过使用来自DNA微阵列杂交实验的基因表达数据来对基因进行功能分类的方法。该方法基于支持向量机(SVM)的理论。 SVM被认为是一种监督的计算机学习方法,因为它们利用了基因功能的先验知识来从表达数据中识别相似函数的未知基因。 SVM避免了与无监督聚类方法相关的几个问题,例如分层聚类和自组织图。 SVM具有许多数学特征,使其对基因表达分析有吸引力,包括它们在选择相似性功能的灵活性,处理大型数据集时解决方案的稀疏性,处理大特征空间的能力以及识别异常值的能力。我们测试了几种使用不同相似性指标以及其他有监督的学习方法的SVM,并发现使用表达数据的SVM最佳识别具有共同函数的基因集。最后,我们使用SVM根据其表达数据来预测未表征的酵母ORF的功能作用。
We introduce a method of functionally classifying genes by using gene expression data from DNA microarray hybridization experiments. The method is based on the theory of support vector machines (SVMs). SVMs are considered a supervised computer learning method because they exploit prior knowledge of gene function to identify unknown genes of similar function from expression data. SVMs avoid several problems associated with unsupervised clustering methods, such as hierarchical clustering and self-organizing maps. SVMs have many mathematical features that make them attractive for gene expression analysis, including their flexibility in choosing a similarity function, sparseness of solution when dealing with large data sets, the ability to handle large feature spaces, and the ability to identify outliers. We test several SVMs that use different similarity metrics, as well as some other supervised learning methods, and find that the SVMs best identify sets of genes with a common function using expression data. Finally, we use SVMs to predict functional roles for uncharacterized yeast ORFs based on their expression data.