Clowder: Open Source Data Management for Long Tail Data

Clowder: Open Source Data Management for Long Tail Data
复制标题

DOI:
10.1145/3219104.3219159
复制
发表时间:
2018-07
期刊:
Proceedings of the Practice and Experience on Advanced Research Computing
影响因子:
--
通讯作者:
Luigi Marini;I. Gutierrez-Polo;R. Kooper;Sandeep Puthanveetil Satheesan;M. Burnette;J. Lee;Todd Nicholson;Yan Zhao;Kenton McHenry
Luigi Marini;I. Gutierrez-Polo;R. Kooper;Sandeep Puthanveetil Satheesan;M. Burnette;J. Lee;Todd Nicholson;Yan Zhao;Kenton McHenry
中科院分区:
其他
文献类型:
--
作者:
Luigi Marini;I. Gutierrez-Polo;R. Kooper;Sandeep Puthanveetil Satheesan;M. Burnette;J. Lee;Todd Nicholson;Yan Zhao;Kenton McHenry

文献摘要

被引文献

相似文献

Clowder是一个开源数据管理系统,支持跨多个研究领域和不同数据类型的长尾数据和元数据的数据管理。机构和实验室可以在本地硬件或远程云计算资源上安装和定制他们自己的框架实例,为分布式研究人员社区提供共享服务。数据可以直接从仪器中获取,也可以由用户手动上传,然后使用web前端与远程合作者共享。我们讨论了在设计和开发一个系统时遇到的一些挑战,这个系统可以很容易地适应不同的科学领域,包括数字保存、地球科学、材料科学、医学、社会科学、文化遗产和艺术。这些挑战包括对大量数据的支持、领域特定预处理算法的横向扩展、在web浏览器中提供新数据可视化的能力、用于自动数据获取和管理的全面web服务API、一套支持用户和算法社区数据注释的社交注释和元数据管理功能,以及与运行在异构集群上的代码交互的基于web的前端。包括高性能计算资源。
Clowder is an open source data management system to support data curation of long tail data and metadata across multiple research domains and diverse data types. Institutions and labs can install and customize their own instance of the framework on local hardware or on remote cloud computing resources to provide a shared service to distributed communities of researchers. Data can be ingested directly from instruments or manually uploaded by users and then shared with remote collaborators using a web front end. We discuss some of the challenges encountered in designing and developing a system that can be easily adapted to different scientific areas including digital preservation, geoscience, material science, medicine, social science, cultural heritage and the arts. Some of these challenges include support for large amounts of data, horizontal scaling of domain specific preprocessing algorithms, ability to provide new data visualizations in the web browser, a comprehensive Web service API for automatic data ingestion and curation, a suite of social annotation and metadata management features to support data annotation by communities of users and algorithms, and a web based front-end to interact with code running on heterogeneous clusters, including HPC resources.