Pattern recognition of gas chromatography mass spectrometry of human volatiles in sweat to distinguish the sex of subjects and determine potential discriminatory marker peaks

Pattern recognition of gas chromatography mass spectrometry of human volatiles in sweat to distinguish the sex of subjects and determine potential discriminatory marker peaks
复制标题

DOI:
10.1016/j.chemolab.2006.12.004
复制
发表时间:
2007-06-15
影响因子:
3.9
通讯作者:
Penn, Dustin J.
Penn, Dustin J.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Dixon, Sarah J.;Xu, Yun;Penn, Dustin J.

文献摘要

被引文献

相似文献

对182名受试者的汗液提取物进行了5次(超过5周)的气相色谱质谱分析,以确定是否有可能将样本分类为男性和女性样本。所有方法均适用于平方根归一化GC-MS峰面积的峰表。使用单变量(t-统计量)和多变量(偏最小二乘判别分析:PLS-DA)方法分别在每两周鉴定潜在的标志物,选择每两周具有高排名的那些峰。使用PLS-DA进行分类,每两周使用100次重复选择模型,将数据随机分为测试集和训练集,并使用自助法找到100个模型中每个模型的显著成分的数量。列联表可以绘制的错误分类的样本的数量,使用三个错误的标准,即自动预测,自举和测试集。可以调整将样本分配到组的决策阈值,并使用接收者操作特征曲线来可视化改变该阈值的影响。它表明,通过使用整个9 - 10测量有一个更密切的对应关系,自动预测和测试集的错误率比182测量,其中有较少的协议,这表明样本量有一个关键的作用。提出了一种研究大型代谢组学数据集的一般策略。(c)2007 Elsevier B. V.保留所有权利。
Pattern recognition studies are performed on the gas chromatography mass spectrometry of extracts of human sweat of 182 subjects sampled 5 times (over 5 fortnights), in an attempt to determine whether it is possible to classify samples into those arising from males and females. All methods were applied to peak tables of square root normalised GC-MS peak areas. Potential markers were identified using both a univariate (t-statistic) and multivariate (Partial Least Squares Discriminant Analysis: PLS-DA) method, on each fortnight separately, selecting those peaks that have high ranks each fortnight. Classification was performed using PLS-DA, selecting the model using 100 repetitions for each fortnight dividing the data into test and training sets randomly, and using the bootstrap to find the number of significant components for each of the 100 models. Contingency tables can be drawn up for the number of misclassified samples, using three error criteria, namely autoprediction, bootstrap and test set. The decision threshold for which sample is assigned to a group can be adjusted and Receiver Operator Characteristic curves were used to visualise the influence on changing this threshold. It is shown that by using the entire set of 9 10 measurements there is a closer correspondence between autoprediction and test set error rates than for 182 measurements where there is less agreement, suggesting that sample size has a key role. A general strategy for studying large metabolomics datasets is proposed. (c) 2007 Elsevier B.V. All rights reserved.