CGAN-based synthetic multivariate time-series generation: a solution to data scarcity in solar flare forecasting

CGAN-based synthetic multivariate time-series generation: a solution to data scarcity in solar flare forecasting
复制标题

DOI:
10.1007/s00521-022-07361-8
复制
发表时间:
2022-05
影响因子:
6
通讯作者:
Yang Chen;Dustin J. Kempton;Azim Ahmadzadeh;Junzhi Wen;Anli Ji;R. Angryk
Yang Chen;Dustin J. Kempton;Azim Ahmadzadeh;Junzhi Wen;Anli Ji;R. Angryk
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yang Chen;Dustin J. Kempton;Azim Ahmadzadeh;Junzhi Wen;Anli Ji;R. Angryk

文献摘要

相似文献

改进监督算法的主要瓶颈之一是数据稀缺。这可能是由许多原因造成的,这些原因往往根源于极其昂贵和冗长的数据收集过程。在太阳物理学等自然领域,可能需要数十年才能获得足够大的样本用于机器学习。受生成对抗网络(GAN)在生成合成图像方面取得的巨大成功的启发,在这项研究中,我们在最近发布的针对太阳耀斑预测的基准数据集上采用了条件GAN(CGAN)。我们的目标是生成合成的多变量时间序列数据,(1)在统计上类似于真实的数据和(2)提高耀斑预测的性能时,用于补救的稀缺性强耀斑。为了评估生成的样本,首先,我们使用Kullback-Leibler散度和对抗性准确性度量来量化真实的和合成数据在描述性统计方面的相似性。其次,我们通过训练预测模型来评估生成的样本对描述性统计的影响,这导致了显着的改善(TSS超过1100%,HSS超过350%)。第三,我们使用生成的时间序列来检查它们对减轻强耀斑稀缺性的高维贡献,与过采样,欠采样和合成过采样方法相比,我们还观察到TSS(4%,7%和31%)和HSS(75%,35%和72%)的显着改善。我们相信我们的发现可以为更强大和准确的耀斑预测模型打开新的大门。
One of the major bottlenecks in refining supervised algorithms is data scarcity. This might be caused by a number of reasons often rooted in extremely expensive and lengthy data collection processes. In natural domains such as Heliophysics, it may take decades for sufficiently large samples for machine learning purposes. Inspired by the massive success of generative adversarial networks (GANs) in generating synthetic images, in this study we employed the conditional GAN (CGAN) on a recently released benchmark dataset tailored for solar flare forecasting. Our goal is to generate synthetic multivariate time-series data that (1) are statistically similar to the real data and (2) improve the performance of flare prediction when used to remedy the scarcity of strong flares. To evaluate the generated samples, first, we used the Kullback–Leibler divergence and adversarial accuracy measures to quantify the similarity between the real and synthetic data in terms of their descriptive statistics. Second, we evaluated the impact of the generated samples by training a predictive model on their descriptive statistics, which resulted in a significant improvement (over 1100% in TSS and 350% in HSS). Third, we used the generated time series to examine their high-dimensional contribution to mitigating the scarcity of the strong flares, which we also observed a significant improvement in terms of TSS (4%, 7%, and 31%) and HSS (75%, 35%, and 72%), compared to oversampling, undersampling, and synthetic oversampling methods, respectively. We believe our findings can open new doors toward more robust and accurate flare forecasting models.