Online scientific data curation, publication, and archiving

Online scientific data curation, publication, and archiving
复制标题

在线科学数据管理、出版和归档

DOI:
10.1117/12.461524
复制
发表时间:
2002
期刊:
--
影响因子:
--
通讯作者:
Jan vandenBerg
Jan vandenBerg
中科院分区:
--
文献类型:
--
作者:
J. Gray;A. Szalay;Anirudha Thakar;C. Stoughton;Jan vandenBerg

文献摘要

被引文献

相似文献

科学项目是数据发布者。当前和未来科学数据的规模和复杂性改变了出版过程的性质。出版物正在成为一个主要的项目组成部分。至少,项目必须保留其收集的临时数据。派生数据可以从元数据重建,但元数据是短暂的。从长远来看,项目应该期望有一些存档来保存数据。我们观察到,已发布的科学数据需要永远可用——这导致了版本的数据金字塔和派生数据量爆炸的数据膨胀。作为示例,本文介绍了斯隆数字巡天 (SDSS) 的数据发布、数据访问、管理和保存策略。
Science projects are data publishers. The scale and complexity of current and future science data changes the nature of the publication process. Publication is becoming a major project component. At a minimum, a project must preserve the ephemeral data it gathers. Derived data can be reconstructed from metadata, but metadata is ephemeral. Longer term, a project should expect some archive to preserve the data. We observe that published scientific data needs to be available forever -- this gives rise to the data pyramid of versions and to data inflation where the derived data volumes explode. As an example, this article describes the Sloan Digital Sky Survey (SDSS) strategies for data publication, data access, curation, and preservation.