UvA-DARE ( Digital Academic Repository ) Centering , scaling and transformations : improving the biological information content of metabolomics data

UvA-DARE ( Digital Academic Repository ) Centering , scaling and transformations : improving the biological information content of metabolomics data
复制标题

DOI:
--
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
R. A. V. D. Berg;H. Hoefsloot;J. Westerhuis;A. Smilde;M. Werf
R. A. V. D. Berg;H. Hoefsloot;J. Westerhuis;A. Smilde;M. Werf
中科院分区:
其他
文献类型:
--
作者:
R. A. V. D. Berg;H. Hoefsloot;J. Westerhuis;A. Smilde;M. Werf

文献摘要

被引文献

相似文献

背景:从大数据集中提取相关生物信息是功能基因组学研究的主要挑战。数据的不同方面阻碍了它们的生物学解释。例如,在代谢组学数据集中,不同代谢物的浓度存在5000倍的差异,而这些差异与这些代谢物的生物学相关性不成比例。然而,数据分析方法无法做出这种区分。数据预处理方法可以通过强调数据集中的生物信息,纠正阻碍代谢组学数据集生物学解释的方面,从而提高其生物学可解释性。结果:在一个真实的代谢组学数据集上测试了不同的数据预处理方法,即定心、自动缩放、帕累托缩放、范围缩放、巨大缩放、对数变换和幂变换。从生物学的角度来看,它们极大地影响了数据分析的结果,从而影响了最重要的代谢物的排名。此外,数据分析前使用的数据预处理方法会影响排序的稳定性、技术误差对数据分析的影响以及数据分析方法对选择高丰代谢物的偏好。结论:不同的预处理方法强调数据的不同方面,每种预处理方法都有各自的优缺点。预处理方法的选择取决于要回答的生物学问题、数据集的性质和所选择的数据分析方法。对于本研究中使用的验证数据集的探索性分析,自动缩放和范围缩放比其他预处理方法效果更好。也就是说,范围缩放和自动缩放能够消除代谢物等级对平均浓度和折叠变化幅度的依赖,并在PCA(主成分分析)后显示出生物学上合理的结果。综上所述,选择合适的数据预处理方法是代谢组学数据分析中必不可少的一步,对鉴定出最重要的代谢物有很大的影响。发表日期:2006年6月8日BMC Genomics 2006,7:142 doi:10.1186/1471-2164-7-142收稿日期:2006年2月20日接收日期:2006年6月8日本文可从:http://www.biomedcentral.com/1471-2164/7/142©2006 van den Berg et al;持牌人中环生物科技有限公司这是一篇在知识共享署名许可(http://creativecommons.org/licenses/by/2.0)下发布的开放获取文章,该许可允许在任何媒体上不受限制地使用、分发和复制,前提是正确引用原始作品。BMC Genomics 2006,7:142 http://www.biomedcentral.com/1471-2164/7/142背景功能基因组学方法越来越多地被用于阐明复杂的生物学问题,其应用范围从人类健康[1]到微生物菌株改善[2]。功能基因组学工具的共同点是,它们旨在测量生物体对感兴趣的环境条件的完整生物分子反应。转录组学和蛋白质组学旨在分别测量所有mRNA和蛋白质,而代谢组学旨在测量所有代谢物[3,4]。在代谢组学研究中,从所研究的生物学状况的采样到对数据分析结果的生物学解释之间有几个步骤(图1)。首先,提取生物样品并准备用于分析。随后,应用不同的数据预处理步骤[3,5],以生成反映(细胞内)代谢物浓度的归一化峰区形式的“干净”数据。这些干净的数据可以作为数据分析的输入。然而,在开始数据分析之前,使用合适的数据预处理方法是很重要的。数据预处理方法将干净的数据转换为不同的尺度(例如,相对尺度或对数尺度)。因此,他们的目标是关注相关(生物)信息,并减少干扰因素(如测量噪声)的影响。可用于数据预处理的程序有缩放、定心和转换。在本文中,我们讨论了代谢组学数据的不同特性,预处理方法如何影响这些特性,以及如何分析数据预处理方法的效果。数据预处理的效果将通过对四种不同碳源上生长的恶臭假单胞菌S12代谢组学数据集的八种数据预处理方法来说明。在代谢组学实验中,获得代谢组的快照,反映了实验条件下所研究的细胞状态或表型。产生本文所用数据集的实验是按照实验设计进行的。在实验设计中,有目的地选择实验条件来诱导感兴趣区域的变化。由此产生的代谢组变异称为诱导生物变异。然而,代谢组学数据中也存在其他因素:1。所测代谢物浓度之间的数量级差异;例如,信号分子的平均浓度远低于像ATP这样含量丰富的化合物的平均浓度。然而,从生物学的角度来看,高浓度代谢物并不一定比低浓度代谢物更重要。2. 诱导变异引起的代谢物浓度的折叠变化的差异;图1生物采样与最重要代谢物排序的步骤不同。生物实验
Background: Extracting relevant biological information from large data sets is a major challenge in functional genomics research. Different aspects of the data hamper their biological interpretation. For instance, 5000-fold differences in concentration for different metabolites are present in a metabolomics data set, while these differences are not proportional to the biological relevance of these metabolites. However, data analysis methods are not able to make this distinction. Data pretreatment methods can correct for aspects that hinder the biological interpretation of metabolomics data sets by emphasizing the biological information in the data set and thus improving their biological interpretability. Results: Different data pretreatment methods, i.e. centering, autoscaling, pareto scaling, range scaling, vast scaling, log transformation, and power transformation, were tested on a real-life metabolomics data set. They were found to greatly affect the outcome of the data analysis and thus the rank of the, from a biological point of view, most important metabolites. Furthermore, the stability of the rank, the influence of technical errors on data analysis, and the preference of data analysis methods for selecting highly abundant metabolites were affected by the data pretreatment method used prior to data analysis. Conclusion: Different pretreatment methods emphasize different aspects of the data and each pretreatment method has its own merits and drawbacks. The choice for a pretreatment method depends on the biological question to be answered, the properties of the data set and the data analysis method selected. For the explorative analysis of the validation data set used in this study, autoscaling and range scaling performed better than the other pretreatment methods. That is, range scaling and autoscaling were able to remove the dependence of the rank of the metabolites on the average concentration and the magnitude of the fold changes and showed biologically sensible results after PCA (principal component analysis). In conclusion, selecting a proper data pretreatment method is an essential step in the analysis of metabolomics data and greatly affects the metabolites that are identified to be the most important. Published: 08 June 2006 BMC Genomics 2006, 7:142 doi:10.1186/1471-2164-7-142 Received: 20 February 2006 Accepted: 08 June 2006 This article is available from: http://www.biomedcentral.com/1471-2164/7/142 © 2006 van den Berg et al; licensee BioMed Central Ltd. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. Page 1 of 15 (page number not for citation purposes) BMC Genomics 2006, 7:142 http://www.biomedcentral.com/1471-2164/7/142 Background Functional genomics approaches are increasingly being used for the elucidation of complex biological questions with applications that range from human health [1] to microbial strain improvement [2]. Functional genomics tools have in common that they aim to measure the complete biomolecule response of an organism to the environmental conditions of interest. While transcriptomics and proteomics aim to measure all mRNA and proteins, respectively, metabolomics aims to measure all metabolites [3,4]. In metabolomics research, there are several steps between the sampling of the biological condition under study and the biological interpretation of the results of the data analysis (Figure 1). First, the biological samples are extracted and prepared for analysis. Subsequently, different data preprocessing steps [3,5] are applied in order to generate 'clean' data in the form of normalized peak areas that reflect the (intracellular) metabolite concentrations. These clean data can be used as the input for data analysis. However, it is important to use an appropriate data pretreatment method before starting data analysis. Data pretreatment methods convert the clean data to a different scale (for instance, relative or logarithmic scale). Hereby, they aim to focus on the relevant (biological) information and to reduce the influence of disturbing factors such as measurement noise. Procedures that can be used for data pretreatment are scaling, centering and transformations. In this paper, we discuss different properties of metabolomics data, how pretreatment methods influence these properties, and how the effects of the data pretreatment methods can be analyzed. The effect of data pretreatment will be illustrated by the application of eight data pretreatment methods to a metabolomics data set of Pseudomonas putida S12 grown on four different carbon sources. Properties of metabolome data In metabolomics experiments, a snapshot of the metabolome is obtained that reflects the cellular state, or phenotype, under the experimental conditions studied [3]. The experiments that resulted in the data set used in this paper were conducted according to an experimental design. In an experimental design, the experimental conditions are purposely chosen to induce variation in the area of interest. The resulting variation in the metabolome is called induced biological variation. However, other factors are also present in metabolomics data: 1. Differences in orders of magnitude between measured metabolite concentrations; for example, the average concentration of a signal molecule is much lower than the average concentration of a highly abundant compound like ATP. However, from a biological point of view, metabolites present in high concentrations are not necessarily more important than those present at low concentrations. 2. Differences in the fold changes in metabolite concentration due to the induced variation; the concentrations of metabolites in the central metabolism are generally relaThe different steps between biological sampling and ranking of the most important m tabolites Figure 1 The different steps between biological sampling and ranking of the most important metabolites. Biological experiment