Toward a Sample Metadata Standard in Public Proteomics Repositories

Toward a Sample Metadata Standard in Public Proteomics Repositories
复制标题

DOI:
10.1021/acs.jproteome.0c00376
复制
发表时间:
2020-10-02
影响因子:
4.4
通讯作者:
Perez-Riverol, Yasset
Perez-Riverol, Yasset
中科院分区:
生物学2区
文献类型:
--
作者:
Perez-Riverol, Yasset

文献摘要

被引文献

相似文献

元数据在蛋白质组学数据库中是必不可少的,对于解释和重新分析存储的数据集至关重要。对于每个蛋白质组学数据集,我们应该捕获至少三个级别的元数据:(i)数据集描述,(ii)样本到数据文件相关信息,以及(iii)标准数据文件格式(例如,mzIdentML、mzML或mzTab)。虽然所有ProteomeXchange合作伙伴都支持数据集描述和标准数据文件格式,但有关样品到数据文件的信息大多缺失。最近,欧洲质谱生物信息学共同体(EuBIC)的成员创建了一个名为蛋白质组学样本到数据文件格式的开源项目(https://github.com/bigbio/proteomics-metadata-standard/),以实现公共蛋白质组学数据集样本元数据的标准化。在这里,该项目被提交给蛋白质组学社区,我们呼吁贡献者,包括研究人员,期刊和财团提供有关格式的反馈。我们相信这项工作将提高可重复性,并促进致力于蛋白质组学数据分析的新工具的开发。
Metadata is essential in proteomics data repositories and is crucial to interpret and reanalyze the deposited data sets. For every proteomics data set, we should capture at least three levels of metadata: (i) data set description, (ii) the sample to data files related information, and (iii) standard data file formats (e.g., mzIdentML, mzML, or mzTab). While the data set description and standard data file formats are supported by all ProteomeXchange partners, the information regarding the sample to data files is mostly missing. Recently, members of the European Bioinformatics Community for Mass Spectrometry (EuBIC) have created an open-source project called Sample to Data file format for Proteomics (https://github.com/bigbio/proteomics-metadata-standard/) to enable the standardization of sample metadata of public proteomics data sets. Here, the project is presented to the proteomics community, and we call for contributors, including researchers, journals, and consortiums to provide feedback about the format. We believe this work will improve reproducibility and facilitate the development of new tools dedicated to proteomics data analysis.