A benchmark spike-in data set for biomarker identification in metabolomics

A benchmark spike-in data set for biomarker identification in metabolomics
复制标题

DOI:
10.1002/cem.1420
复制
发表时间:
2012-01-01
影响因子:
2.4
通讯作者:
Wehrens, Ron
Wehrens, Ron
中科院分区:
化学3区
文献类型:
--
作者:
Franceschi, Pietro;Masuero, Domenico;Wehrens, Ron

文献摘要

被引文献

相似文献

生物标志物选择的创新方法的开发和验证在许多组学技术中至关重要。不幸的是,在真实的数据上实际测试新方法是困难的,因为在真实的数据集中,人们永远无法确定真正的生物标志物。在本文中,我们提出了一个公开的代谢组学超高效液相色谱质谱加标数据集的苹果。数据集由10个对照样品和3个相同大小的加标组组成,其中天然存在的化合物以不同浓度添加。在这个意义上说,数据集可以作为一个测试床,以评估新算法的性能,并将它们与以前发表的结果进行比较,我们说明了一些可能性,通过比较两种流行的生物标记选择方法,单变量t检验和多变量的重要性投影的性能,这穗在数据集。为了促进数据的广泛使用,提供了原始数据文件以及预处理的峰列表。版权所有(C)2012约翰威利父子有限公司
The development and the validation of innovative approaches for biomarker selection are of paramount importance in many -omics technologies. Unfortunately, the actual testing of new methods on real data is difficult, because in real data sets, one can never be sure about the true biomarkers. In this paper, we present a publicly available metabolomic ultra performance liquid chromatographymass spectrometry spike-in data set for apples. The data set consists of 10 control samples and three spiked sets of the same size, where naturally occurring compounds are added in different concentrations. In this sense, the data set can serve as a test bed to assess the performance of new algorithms and compare them with previously published results.We illustrate some of the possibilities provided by this spike-in data set by comparing the performance of two popular biomarker-selection methods, the univariate t-test and the multivariate variable importance in projection. To promote a widespread use of the data, raw data files as well as preprocessed peak lists are made available. Copyright (C) 2012 John Wiley & Sons, Ltd.