The Materials Data Facility: Data Services to Advance Materials Science Research

The Materials Data Facility: Data Services to Advance Materials Science Research
复制标题

DOI:
10.1007/s11837-016-2001-3
复制
发表时间:
2016-08-01
期刊:
JOM
影响因子:
2.6
通讯作者:
Foster, I.
Foster, I.
中科院分区:
材料科学3区
文献类型:
--
作者:
Blaiszik, B.;Chard, K.;Foster, I.

文献摘要

被引文献

相似文献

随着资助机构和机构对数据管理的要求越来越严格,对研究可复制性挑战的关注不断扩大,以及数据规模和异质性的不断增长,材料界正在出现新的数据需求。材料数据设施(Materials Data Facility,简称MADF)运营着两项云托管服务,即数据发布和数据发现,其功能包括促进开放式数据共享、自助式数据发布和管理,并鼓励数据重用,以及强大的数据发现工具。数据发布服务简化了将数据复制到安全存储位置、为数据分配可引用的持久标识符以及记录自定义(例如,材料、技术或仪器特定的)和自动提取的元数据,而数据发现服务将提供高级搜索能力(例如,分面、自由文本范围查询和全文搜索)。该服务使个人研究人员,研究项目和机构能够(I)从本地存储,机构数据存储或云存储发布研究数据集,无论大小,而无需第三方发布者的参与;(II)构建,共享和实施可扩展的特定于域的自定义元数据模式;(III)通过代表性状态转移(REST)应用程序接口(API)与发布的数据和元数据交互,以促进自动化、分析和反馈;以及(IV)访问数据发现模型,该模型允许研究人员搜索、询问并最终建立在现有的已发布数据上。我们描述了该公司的设计、现状和未来计划。
With increasingly strict data management requirements from funding agencies and institutions, expanding focus on the challenges of research replicability, and growing data sizes and heterogeneity, new data needs are emerging in the materials community. The materials data facility (MDF) operates two cloud-hosted services, data publication and data discovery, with features to promote open data sharing, self-service data publication and curation, and encourage data reuse, layered with powerful data discovery tools. The data publication service simplifies the process of copying data to a secure storage location, assigning data a citable persistent identifier, and recording custom (e.g., material, technique, or instrument specific) and automatically-extracted metadata in a registry while the data discovery service will provide advanced search capabilities (e.g., faceting, free text range querying, and full text search) against the registered data and metadata. The MDF services empower individual researchers, research projects, and institutions to (I) publish research datasets, regardless of size, from local storage, institutional data stores, or cloud storage, without involvement of third-party publishers; (II) build, share, and enforce extensible domain-specific custom metadata schemas; (III) interact with published data and metadata via representational state transfer (REST) application program interfaces (APIs) to facilitate automation, analysis, and feedback; and (IV) access a data discovery model that allows researchers to search, interrogate, and eventually build on existing published data. We describe MDF's design, current status, and future plans.