STanford EArthquake Dataset (STEAD): A Global Data Set of Seismic Signals for AI

STanford EArthquake Dataset (STEAD): A Global Data Set of Seismic Signals for AI
复制标题

DOI:
10.1109/access.2019.2947848
复制
发表时间:
2019-01-01
期刊:
影响因子:
3.9
通讯作者:
Beroza, Gregory C.
Beroza, Gregory C.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Mousavi, S. Mostafa;Sheng, Yixiao;Beroza, Gregory C.

文献摘要

被引文献

相似文献

地震学是一门数据丰富、数据驱动的科学。应用机器学习从地震数据中获得新的见解是地震学的一个快速发展的子领域。大量地震数据和计算资源的可用性,以及先进技术的发展,可以建立更强大的模型和算法来处理和分析地震信号。已知的例子或标记的数据集,是建立监督模型的必要条件。地震学已经有了标记数据,但这些标记的可靠性变化很大,缺乏高质量的标记数据集作为基础事实,以及缺乏标准基准,这些都是取得更快进展的障碍。在本文中,我们提出了一个高质量的、大尺度的、由地震仪器记录的局部地震和非地震信号的全球数据集。目前状态下的数据集包含两类:(1)局部地震波形(在距离地震350公里以内的“局部”距离记录)和(2)没有地震信号的地震噪声波形。这些数据共同构成了120万个时间序列或超过19000小时的地震信号记录。构建这样一个具有可靠标签的大规模数据库是一项具有挑战性的任务。在这里,我们介绍了数据集的属性,描述了数据收集,质量控制程序,以及我们为确保准确标记所采取的处理步骤,并讨论了潜在的应用。我们希望STEAD的规模和准确性为地震学界和其他领域的研究人员提供新的和无与伦比的机会。
Seismology is a data rich and data-driven science. Application of machine learning for gaining new insights from seismic data is a rapidly evolving sub-field of seismology. The availability of a large amount of seismic data and computational resources, together with the development of advanced techniques can foster more robust models and algorithms to process and analyze seismic signals. Known examples or labeled data sets, are the essential requisite for building supervised models. Seismology has labeled data, but the reliability of those labels is highly variable, and the lack of high-quality labeled data sets to serve as ground truth as well as the lack of standard benchmarks are obstacles to more rapid progress. In this paper we present a high-quality, large-scale, and global data set of local earthquake and non-earthquake signals recorded by seismic instruments. The data set in its current state contains two categories: (1) local earthquake waveforms (recorded at "local" distances within 350 km of earthquakes) and (2) seismic noise waveforms that are free of earthquake signals. Together these data comprise similar to 1.2 million time series or more than 19,000 hours of seismic signal recordings. Constructing such a large-scale database with reliable labels is a challenging task. Here, we present the properties of the data set, describe the data collection, quality control procedures, and processing steps we undertook to insure accurate labeling, and discuss potential applications. We hope that the scale and accuracy of STEAD presents new and unparalleled opportunities to researchers in the seismological community and beyond.