AstroCatR: a mechanism and tool for efficient time series reconstruction of large-scale astronomical catalogues

AstroCatR: a mechanism and tool for efficient time series reconstruction of large-scale astronomical catalogues
复制标题

AstroCatR:大规模天文目录高效时间序列重建的机制和工具

DOI:
10.1093/mnras/staa1413
复制
发表时间:
2020
影响因子:
4.8
通讯作者:
Zhao Qing
Zhao Qing
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Yu Ce;Li Kun;Tang Shanjiang;Sun Chao;Ma Bin;Zhao Qing

文献摘要

被引文献

相似文献

在时间域天文学中,天体的时间序列数据通常被用来研究太阳系外行星、超新星等有价值的、意外的天体。由于数据量的快速增长,传统的人工方法对于连续分析积累的观测数据变得极其困难和不可行。为了满足这种需求,我们设计并实现了一个专门的工具AstroCatR,它可以高效、灵活地从大规模天文星表中重建时间序列数据。AstroCatR可以从灵活图像传输系统(FITS)文件或数据库加载原始目录数据,匹配每个项目以确定它属于哪个对象,并最终生成时间序列数据集。为了支持大规模数据集的高性能并行处理,AstroCatR使用提取-转换-加载(ETL)预处理模块来创建天空区域文件并平衡工作负载。匹配模块使用重叠索引法和内存中的参考表来提高准确率和性能。AstroCatR的输出可以存储在CSV文件中,也可以根据需要转换为其他格式。同时,基于模块的软件架构确保了AstroCatR的灵活性和可扩展性。我们使用三个南极调查望远镜(AST3)的实际观测数据来评估AstroCatR。实验表明,通过设置相关参数和配置文件,AstroCatR可以高效、灵活地重建所有时间序列数据。此外,在匹配海量目录时,该工具比使用关系数据库管理系统的方法快约3倍。
Time series data of celestial objects are commonly used to study valuable and unexpected objects such as extrasolar planets and supernova in time domain astronomy. Due to the rapid growth of data volume, traditional manual methods are becoming extremely hard and infeasible for continuously analysing accumulated observation data. To meet such demands, we designed and implemented a special tool named AstroCatR that can efficiently and flexibly reconstruct time series data from large-scale astronomical catalogues. AstroCatR can load original catalogue data from Flexible Image Transport System (FITS) files or data bases, match each item to determine which object it belongs to, and finally produce time series data sets. To support the high-performance parallel processing of large-scale data sets, AstroCatR uses the extract-transform-load (ETL) pre-processing module to create sky zone files and balance the workload. The matching module uses the overlapped indexing method and an in-memory reference table to improve accuracy and performance. The output of AstroCatR can be stored in CSV files or be transformed other into formats as needed. Simultaneously, the module-based software architecture ensures the flexibility and scalability of AstroCatR. We evaluated AstroCatR with actual observation data from The three Antarctic Survey Telescopes (AST3). The experiments demonstrate that AstroCatR can efficiently and flexibly reconstruct all time series data by setting relevant parameters and configuration files. Furthermore, the tool is approximately 3× faster than methods using relational data base management systems at matching massive catalogues.