MIQS: metadata indexing and querying service for self-describing file formats

MIQS: metadata indexing and querying service for self-describing file formats
复制标题

DOI:
10.1145/3295500.3356146
复制
发表时间:
2019-11
期刊:
Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Wei Zhang;S. Byna;Houjun Tang;Brody Williams;Yong Chen
Wei Zhang;S. Byna;Houjun Tang;Brody Williams;Yong Chen
中科院分区:
其他
文献类型:
--
作者:
Wei Zhang;S. Byna;Houjun Tang;Brody Williams;Yong Chen

文献摘要

被引文献

相似文献

科学应用程序通常以自描述数据文件格式存储数据集,例如HDF5和netCDF。遗憾的是,由于数据集的庞大规模,有效地搜索这些文件中的元数据仍然具有挑战性。现有的解决方案提取元数据并将其存储在外部数据库管理系统(DBMS)中以定位所需的数据。然而,这种做法在提取和查询中引入了显著的开销和复杂性。在这项研究中,我们提出了一种新的元数据索引和查询服务(MIQS),它删除了外部数据库管理系统,并利用内存索引来实现高效的元数据搜索。MIQS遵循自包含的数据管理范例,并为自描述文件格式提供可移植的和无模式的元数据索引和查询功能。我们已经使用最先进的基于MongoDB的元数据索引解决方案评估了MIQS。MIQS在索引构建方面实现了高达99%的时间减少,搜索性能提高了172kx,内存占用减少了75%。
Scientific applications often store datasets in self-describing data file formats, such as HDF5 and netCDF. Regrettably, to efficiently search the metadata within these files remains challenging due to the sheer size of the datasets. Existing solutions extract the metadata and store it in external database management systems (DBMS) to locate desired data. However, this practice introduces significant overhead and complexity in extraction and querying. In this research, we propose a novel Metadata Indexing and Querying Service (MIQS), which removes the external DBMS and utilizes in-memory index to achieve efficient metadata searching. MIQS follows the self-contained data management paradigm and provides portable and schema-free metadata indexing and querying functionalities for self-describing file formats. We have evaluated MIQS with the state-of-the-art MongoDB-based metadata indexing solution. MIQS achieved up to 99% time reduction in index construction and up to 172kx search performance improvement with up to 75% reduction in memory footprint.