Extraction, interpretation and validation of information for comparing samples in metabolic LC/MS data sets

Extraction, interpretation and validation of information for comparing samples in metabolic LC/MS data sets
复制标题

DOI:
10.1039/b501890k
复制
发表时间:
2005-01-01
期刊:
影响因子:
4.2
通讯作者:
Antti, H
Antti, H
中科院分区:
化学2区
文献类型:
--
作者:
Jonsson, P;Bruce, SJ;Antti, H

文献摘要

被引文献

相似文献

LC/MS是一种分析技术,由于其高灵敏度,已经变得越来越流行,用于生成生物样品中的代谢特征和用于建立代谢数据库。然而,为了能够创建强大且可解释的(透明)多变量模型用于比较许多样本,数据必须满足某些特定标准:(i)每个样本的特征在于相同数量的变量,(ii)这些变量中的每一个都在所有观测中表示,和(iii)一个样品中的变量具有相同的生物学意义或在所有其他样品中代表相同的代谢物。此外,获得的模型必须有能力作出预测,e。G.相关和独立的样本,其特征在于相应的模型样本。该方法涉及代表性数据集的构建,包括自动峰检测、对齐、保留时间窗口设置、色谱维度求和和通过交替回归进行的数据压缩,其中保留相关代谢变化以用于使用多变量分析进行进一步建模。这种方法的优点是允许基于LC/MS代谢谱比较大量样品,而且还可以创建用于解释所研究的生物系统的方法。这包括在样本中找到相关的系统模式,识别有影响力的变量,验证原始数据中的发现,最后使用模型进行预测。本文将所提出的策略应用于使用来自两个队列(中华人民共和国山西和檀香山(美国))的尿液样本的人群研究。结果表明,使用偏最小二乘判别分析(PLS-DA)提取的信息数据的评估提供了一个强大的,预测和透明的模型,两个群体之间的代谢差异。提出的研究结果表明,这是一个通用的方法,数据处理,分析和大型代谢LC/MS数据集的评价。
LC/MS is an analytical technique that, due to its high sensitivity, has become increasingly popular for the generation of metabolic signatures in biological samples and for the building of metabolic data bases. However, to be able to create robust and interpretable ( transparent) multivariate models for the comparison of many samples, the data must fulfil certain specific criteria: (i) that each sample is characterized by the same number of variables, (ii) that each of these variables is represented across all observations, and (iii) that a variable in one sample has the same biological meaning or represents the same metabolite in all other samples. In addition, the obtained models must have the ability to make predictions of, e. g. related and independent samples characterized accordingly to the model samples. This method involves the construction of a representative data set, including automatic peak detection, alignment, setting of retention time windows, summing in the chromatographic dimension and data compression by means of alternating regression, where the relevant metabolic variation is retained for further modelling using multivariate analysis. This approach has the advantage of allowing the comparison of large numbers of samples based on their LC/MS metabolic profiles, but also of creating a means for the interpretation of the investigated biological system. This includes finding relevant systematic patterns among samples, identifying influential variables, verifying the findings in the raw data, and finally using the models for predictions. The presented strategy was here applied to a population study using urine samples from two cohorts, Shanxi (People's Republic of China) and Honolulu ( USA). The results showed that the evaluation of the extracted information data using partial least square discriminant analysis (PLS-DA) provided a robust, predictive and transparent model for the metabolic differences between the two populations. The presented findings suggest that this is a general approach for data handling, analysis, and evaluation of large metabolic LC/MS data sets.