Data standards can boost metabolomics research, and if there is a will, there is a way.

Data standards can boost metabolomics research, and if there is a will, there is a way.
复制标题

DOI:
10.1007/s11306-015-0879-3
复制
发表时间:
2016
期刊:
Metabolomics : Official journal of the Metabolomic Society
影响因子:
--
通讯作者:
Neumann S
Neumann S
中科院分区:
其他
文献类型:
--
作者:
Rocca-Serra P;Salek RM;Arita M;Correa E;Dayalan S;Gonzalez-Beltran A;Ebbels T;Goodacre R;Hastings J;Haug K;Koulman A;Nikolski M;Oresic M;Sansone SA;Schober D;Smith J;Steinbeck C;Viant MR;Neumann S

文献摘要

被引文献

相似文献

每年都有数千篇使用代谢组学方法的文章发表。随着产生的数据越来越多,仅仅将研究描述为手稿中的文本已经不足以实现再利用:基础数据需要与文献中的发现一起发表,以最大限度地提高公共和私人支出的效益,并利用巨大的机会提高代谢组学和相关学科的科学再现性。关于代谢组学的报告建议大约在十年前开始出现,主要涉及必须在文献中报告的信息清单,以保持一致性。近年来,代谢组学数据标准得到了广泛的发展,包括主要的研究数据,衍生的结果和实验描述,重要的是以机器可读的方式的元数据。这包括独立于供应商的数据标准,例如用于质谱分析的mzML和用于NMR原始数据的nmrML,它们都使科学界能够开发先进的数据处理算法。ISA-Tab等标准涵盖了基本的元数据,包括实验设计、应用的方案、样品之间的关联、数据文件和用于进一步统计分析的实验因素。总之,它们为可重复研究和数据重用(包括元分析)铺平了道路。准备符合标准的数据集的进一步激励措施包括发布数据集的新机会,但也需要在科学期刊的作者指南中进行一点“施压”,以将数据集提交给公共存储库,如NIH代谢组学中心或EMBL-EBI的代谢光。在本文中,我们将研究数据共享的标准,调查它们对代谢组学的影响,并提出建议以提高它们的采用率。
Thousands of articles using metabolomics approaches are published every year. With the increasing amounts of data being produced, mere description of investigations as text in manuscripts is not sufficient to enable re-use anymore: the underlying data needs to be published together with the findings in the literature to maximise the benefit from public and private expenditure and to take advantage of an enormous opportunity to improve scientific reproducibility in metabolomics and cognate disciplines. Reporting recommendations in metabolomics started to emerge about a decade ago and were mostly concerned with inventories of the information that had to be reported in the literature for consistency. In recent years, metabolomics data standards have developed extensively, to include the primary research data, derived results and the experimental description and importantly the metadata in a machine-readable way. This includes vendor independent data standards such as mzML for mass spectrometry and nmrML for NMR raw data that have both enabled the development of advanced data processing algorithms by the scientific community. Standards such as ISA-Tab cover essential metadata, including the experimental design, the applied protocols, association between samples, data files and the experimental factors for further statistical analysis. Altogether, they pave the way for both reproducible research and data reuse, including meta-analyses. Further incentives to prepare standards compliant data sets include new opportunities to publish data sets, but also require a little “arm twisting” in the author guidelines of scientific journals to submit the data sets to public repositories such as the NIH Metabolomics Workbench or MetaboLights at EMBL-EBI. In the present article, we look at standards for data sharing, investigate their impact in metabolomics and give suggestions to improve their adoption.