Advanced Cyberinfrastructure to Enable Search of Big Climate Datasets in THREDDS

Advanced Cyberinfrastructure to Enable Search of Big Climate Datasets in THREDDS
复制标题

DOI:
10.3390/ijgi8110494
复制
发表时间:
2019-11
期刊:
ISPRS Int. J. Geo Inf.
影响因子:
--
通讯作者:
Juozas Gaigalas;L. Di;Ziheng Sun
Juozas Gaigalas;L. Di;Ziheng Sun
中科院分区:
其他
文献类型:
--
作者:
Juozas Gaigalas;L. Di;Ziheng Sun

文献摘要

相似文献

了解气候的过去、现在和变化行为需要来自许多科学领域的大量研究人员的密切合作。目前,由于数据量的急剧增加,在发现、共享和整合气候数据方面存在困难,这极大地限制了必要的跨学科合作。本文讨论了解决异构地球系统观测与建模(ESOM)数据的传输、处理和服务元数据时遇到的相关问题的方法和技术。提出了一种基于网络基础设施的解决方案,通过利用最先进的Web服务技术和爬行现有的数据中心,实现对大型气候数据集的有效编目和两步搜索。为了验证其可行性,UCAR THREDDS数据服务器(TDS)提供PB级ESOM数据,每天更新数百TB的数据,作为案例研究数据集。设计了一个完整的工作流程来分析TDS中的元数据结构并为数据参数创建索引。一个简化的注册模型,定义恒定的信息,划定次要信息,并利用空间和时间的元数据的一致性。该模型推导出一种高性能并发网络爬虫机器人的采样策略,该策略用于在不占用大量网络和计算资源的情况下镜像大数据存档的基本元数据。元数据模型、爬虫和符合标准的目录服务形成了一个增量搜索网络基础设施,使科学家能够近实时地搜索大型气候数据集。该方法在UCAR TDS上进行了测试,结果表明,该方法实现了其设计目标,至少将爬行速度提高了10倍,并将冗余元数据从1.85 GB减少到2.2 MB,这是使当前大多数不可搜索的气候数据服务器可搜索的重大突破。
Understanding the past, present, and changing behavior of the climate requires close collaboration of a large number of researchers from many scientific domains. At present, the necessary interdisciplinary collaboration is greatly limited by the difficulties in discovering, sharing, and integrating climatic data due to the tremendously increasing data size. This paper discusses the methods and techniques for solving the inter-related problems encountered when transmitting, processing, and serving metadata for heterogeneous Earth System Observation and Modeling (ESOM) data. A cyberinfrastructure-based solution is proposed to enable effective cataloging and two-step search on big climatic datasets by leveraging state-of-the-art web service technologies and crawling the existing data centers. To validate its feasibility, the big dataset served by UCAR THREDDS Data Server (TDS), which provides Petabyte-level ESOM data and updates hundreds of terabytes of data every day, is used as the case study dataset. A complete workflow is designed to analyze the metadata structure in TDS and create an index for data parameters. A simplified registration model which defines constant information, delimits secondary information, and exploits spatial and temporal coherence in metadata is constructed. The model derives a sampling strategy for a high-performance concurrent web crawler bot which is used to mirror the essential metadata of the big data archive without overwhelming network and computing resources. The metadata model, crawler, and standard-compliant catalog service form an incremental search cyberinfrastructure, allowing scientists to search the big climatic datasets in near real-time. The proposed approach has been tested on UCAR TDS and the results prove that it achieves its design goal by at least boosting the crawling speed by 10 times and reducing the redundant metadata from 1.85 gigabytes to 2.2 megabytes, which is a significant breakthrough for making the current most non-searchable climate data servers searchable.