Analysis of metabolomic data using support vector machines

Analysis of metabolomic data using support vector machines
复制标题

DOI:
10.1021/ac800954c
复制
发表时间:
2008-10-01
影响因子:
7.4
通讯作者:
Slupsky, Carolyn M.
Slupsky, Carolyn M.
中科院分区:
化学1区
文献类型:
--
作者:
Mahadevan, Sankar;Shah, Sirish L.;Slupsky, Carolyn M.

文献摘要

被引文献

相似文献

代谢组学是一个新兴的领域提供洞察生理过程。通过观察各种生物流体中代谢物浓度的变化,它是研究疾病诊断或进行毒理学研究的有效工具。多变量统计分析通常与核磁共振(NMR)或质谱(MS)数据一起使用,以确定组之间的差异(例如患病与健康)。特征预测模型可以基于一组训练数据来构建,并且这些模型随后用于预测新的测试数据福尔斯是否落入特定类别。在本研究中,通过对从健康受试者(男性和女性)和患有肺炎链球菌的患者获得的尿样进行H-1 NMR光谱来获得代谢组学数据。我们比较了传统PLS-DA多变量分析与支持向量机(SVM)的性能,支持向量机是一种广泛用于基因组研究的技术,用于两个案例研究:(1)可以看到几乎完全区别的案例(健康与肺炎)和(2)区别更模糊的案例(男性与女性)。我们表明,支持向量机是上级PLS-DA在这两种情况下,在预测精度与最少的功能。与PLS-DA相比,支持向量机具有更少的特征,能够提供更好的预测模型。
Metabolomics is an emerging field providing insight into physiological processes. It is an effective tool to investigate disease diagnosis or conduct toxicological studies by observing changes in metabolite concentrations in various biofluids. Multivariate statistical analysis is generally employed with nuclear magnetic resonance (NMR) or mass spectrometry (MS) data to determine differences between groups (for instance diseased vs healthy). Characteristic predictive models may be built based on a set of training data, and these models are subsequently used to predict whether new test data falls under a specific class. In this study, metabolomic data is obtained by doing a H-1 NMR spectroscopy on urine samples obtained from healthy subjects (male and female) and patients suffering from Streptococcus pneumoniae. We compare the performance of traditional PLS-DA multivariate analysis to support vector machines (SVMs), a technique widely used in genome studies on two case studies: (1) a case where nearly complete distinction may be seen (healthy versus pneumonia) and (2) a case where distinction is more ambiguous (male versus female). We show that SVMs are superior to PLS-DA in both cases in terms of predictive accuracy with the least number of features. With fewer number of features, SVMs are able to give better predictive model when compared to that of PLS-DA.