ccSVM: correcting Support Vector Machines for confounding factors in biological data classification.

ccSVM: correcting Support Vector Machines for confounding factors in biological data classification.
复制标题

DOI:
10.1093/bioinformatics/btr204
复制
发表时间:
2011-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Borgwardt K
Borgwardt K
中科院分区:
其他
文献类型:
--
作者:
Li L;Rakitsch B;Borgwardt K

文献摘要

参考文献

被引文献

相似文献

动机:将生物数据分类为不同的组是生物信息学的核心任务:例如,预测基因或蛋白质的功能、患者的疾病状态或基于其基因型的个体表型。支持向量机是一种广泛使用的生物数据分类方法,因为它们具有很高的准确性,处理结构化数据(如字符串)的能力,并且易于集成各种类型的数据。然而,如何在支持向量机分类中纠正混杂因素,如人口结构、年龄或性别或实验条件,目前还不清楚。结果:在本文中,我们提出了一个支持向量机分类器,可以对观察到的混杂因素的预测进行校正。这是通过最小化分类器和混杂因素之间的统计依赖性来实现的。我们证明了这个公式可以转换成一个标准的支持向量机与重新缩放的输入数据。在我们的实验中,我们的混杂校正支持向量机(ccSVM)改进了基于不同实验室样本的肿瘤诊断,不同年龄,种族和性别患者的结核病诊断,以及存在人群结构的表型预测,并且在预测精度方面优于最先进的方法。可用性:ccsvm在MATLAB中的实现可从http://webdav.tuebingen.mpg.de/u/karsten/Forschung/ISMB11_ccSVM/获得。联系人:limin.li@tuebingen.mpg.de;karsten.borgwardt@tuebingen.mpg.de
Motivation: Classifying biological data into different groups is a central task of bioinformatics: for instance, to predict the function of a gene or protein, the disease state of a patient or the phenotype of an individual based on its genotype. Support Vector Machines are a wide spread approach for classifying biological data, due to their high accuracy, their ability to deal with structured data such as strings, and the ease to integrate various types of data. However, it is unclear how to correct for confounding factors such as population structure, age or gender or experimental conditions in Support Vector Machine classification. Results: In this article, we present a Support Vector Machine classifier that can correct the prediction for observed confounding factors. This is achieved by minimizing the statistical dependence between the classifier and the confounding factors. We prove that this formulation can be transformed into a standard Support Vector Machine with rescaled input data. In our experiments, our confounder correcting SVM (ccSVM) improves tumor diagnosis based on samples from different labs, tuberculosis diagnosis in patients of varying age, ethnicity and gender, and phenotype prediction in the presence of population structure and outperforms state-of-the-art methods in terms of prediction accuracy. Availability: A ccSVM-implementation in MATLAB is available from http://webdav.tuebingen.mpg.de/u/karsten/Forschung/ISMB11_ccSVM/. Contact: limin.li@tuebingen.mpg.de; karsten.borgwardt@tuebingen.mpg.de
DOI: 10.1186/1471-2105-6-265
发表时间: 2005-11-04
期刊: BMC bioinformatics
影响因子: 3
作者:
Warnat P;Eils R;Brors B
通讯作者: Brors B
DOI: 10.1038/nature09247
发表时间: 2010-08-19
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1093/bioinformatics/btl386
发表时间: 2006-10-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cawley, Gavin C.;Talbot, Nicola L. C.
通讯作者: Talbot, Nicola L. C.
DOI: 10.1093/bioinformatics/bth294
发表时间: 2004-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Lanckriet, GRG;De Bie, T;Noble, WS
通讯作者: Noble, WS
DOI: 10.1038/nbt1206-1565
发表时间: 2006-12-01
影响因子: 46.9
作者:
Noble, William S.
通讯作者: Noble, William S.