Variable selection using iterative reformulation of training set models for discrimination of samples: application to gas chromatography/mass spectrometry of mouse urinary metabolites.

Variable selection using iterative reformulation of training set models for discrimination of samples: application to gas chromatography/mass spectrometry of mouse urinary metabolites.
复制标题

DOI:
10.1021/ac900251c
复制
发表时间:
2009-07-01
影响因子:
7.4
通讯作者:
Brereton, Richard G.
Brereton, Richard G.
中科院分区:
化学1区
文献类型:
--
作者:
Wongravee, Kanet;Heinrich, Nina;Holmboe, Maria;Schaefer, Michele L.;Reed, Randall R.;Trevejo, Jose;Brereton, Richard G.

文献摘要

参考文献

被引文献

相似文献

本文讨论了变量选择在大代谢组学研究中的应用,并以441只小鼠的尿气相色谱为例,在三个实验中检测年龄、饮食和应激对其化学信号的影响。应用偏最小二乘判别分析(PLS-DA)来获得类别模型,使用20,000次迭代的程序,包括用于模型优化的自举和随机分割成用于验证的测试和训练集。使用训练集上的PLS回归系数选择变量,该训练集使用从自举获得的优化数量的分量。变量按显著性顺序排列,总体最佳变量被选为在100个不同的测试和训练集分割中表现出高度显著性的变量。成本效益分析的变量数量减少的模型进行说明。本文提供了一种策略,适当验证的方法,以确定哪些变量是最显着的两组之间的大型代谢数据集的区别,避免过拟合的常见陷阱,如果变量被选择在一个组合的训练和测试集,并考虑到不同的变量可以选择每次使用迭代程序的样本被分成训练和测试集。
The paper discusses variable selection as used in large metabolomic studies, exemplified by mouse urinary gas chromatography of 441 mice in three experiments to detect the influence of age, diet and stress on their chemosignal. Partial Least Squares Discriminant Analysis (PLS-DA) was applied to obtain class models, using a procedure of 20,000 iterations including the bootstrap for model optimisation and random splits into test and training sets for validation. Variables are selected using PLS regression coefficients on the training set using an optimised number of components obtained from the bootstrap. The variables are ranked in order of significance and the overall optimal variables are selected as those that appear as highly significant over 100 different test and training set splits. Cost benefit analysis of performing the model on a reduced number of variables is also illustrated. This paper provides a strategy for properly validated methods for determining which variables are most significant for discriminating between two groups in large metabolomic datasets avoiding the common pitfall of overfitting if variables are selected on a combined training and test set, and also taking into account that different variables may be selected each time the samples are split into training and test sets using iterative procedures.
DOI: 10.1002/cem.1189
发表时间: 2009-01-01
影响因子: 2.4
作者:
Dixon, Sarah J.;Heinrich, Nina;Brereton, Richard G.
通讯作者: Brereton, Richard G.
DOI: 10.1093/bioinformatics/bti102
发表时间: 2005-04-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Jarvis, RM;Goodacre, R
通讯作者: Goodacre, R
DOI: 10.1016/j.chemolab.2005.05.004
发表时间: 2006-01-20
影响因子: 3.9
作者:
Brown, CD;Davis, HT
通讯作者: Davis, HT
DOI: 10.1016/j.chemolab.2004.09.007
发表时间: 2005-03-28
影响因子: 3.9
作者:
Lima, SLT;Mello, C;Poppi, RJ
通讯作者: Poppi, RJ
DOI: 10.1016/j.chemolab.2006.12.004
发表时间: 2007-06-15
影响因子: 3.9
作者:
Dixon, Sarah J.;Xu, Yun;Penn, Dustin J.
通讯作者: Penn, Dustin J.