Cycle-StarNet: Bridging the Gap between Theory and Data by Leveraging Large Data Sets

Cycle-StarNet: Bridging the Gap between Theory and Data by Leveraging Large Data Sets
复制标题

DOI:
10.3847/1538-4357/abca96
复制
发表时间:
2020-07
期刊:
The Astrophysical Journal
影响因子:
--
通讯作者:
T. O'Briain;Y. Ting 丁;S. Fabbro;K. M. Yi;K. Venn;Spencer Bialek
T. O'Briain;Y. Ting 丁;S. Fabbro;K. M. Yi;K. Venn;Spencer Bialek
中科院分区:
其他
文献类型:
--
作者:
T. O'Briain;Y. Ting 丁;S. Fabbro;K. M. Yi;K. Venn;Spencer Bialek

文献摘要

被引文献

相似文献

恒星光谱数据采集的进步使得有必要在有效的数据分析技术方面实现类似的改进。当前用于分析光谱的自动化方法是(a)数据驱动的,其需要恒星参数和元素丰度的先验知识,或者(B)基于理论合成模型,其易受理论与实践之间的差距的影响。在这项研究中,我们提出了一种混合生成域自适应方法,通过将无监督学习应用于大型光谱调查,将模拟恒星光谱转化为现实光谱。我们将我们的技术应用于R = 22,500的APOGEE H带光谱和Kurucz合成模型。作为概念的证明,两个案例研究。第一个是校准合成数据,使之与观测数据相一致。为了实现这一点,合成模型被变形为类似于观测的光谱,从而减少理论和观测之间的差距。拟合观测到的光谱表明,改进的平均值从1.97降低到1.22,沿着,归一化通量的平均残差从0.16降低到-0.01。第二个案例研究是识别合成建模中缺失谱线的元素来源。一个模拟数据集是用来表明,吸收线可以恢复时,他们是不存在的域之一。该方法可以应用于使用大数据集的其他领域,并且目前受到建模精度的限制。本研究中使用的代码在GitHub上公开提供(https://github.com/teaghan/Cycle_SN)。
Advancements in stellar spectroscopy data acquisition have made it necessary to accomplish similar improvements in efficient data analysis techniques. Current automated methods for analyzing spectra are either (a) data driven, which requires prior knowledge of stellar parameters and elemental abundances, or (b) based on theoretical synthetic models that are susceptible to the gap between theory and practice. In this study, we present a hybrid generative domain-adaptation method that turns simulated stellar spectra into realistic spectra by applying unsupervised learning to large spectroscopic surveys. We apply our technique to the APOGEE H-band spectra at R = 22,500 and the Kurucz synthetic models. As a proof of concept, two case studies are presented. The first is the calibration of synthetic data to become consistent with observations. To accomplish this, synthetic models are morphed into spectra that resemble observations, thereby reducing the gap between theory and observations. Fitting the observed spectra shows an improved average reduced from 1.97 to 1.22, along with a mean residual reduced from 0.16 to −0.01 in normalized flux. The second case study is the identification of the elemental source of missing spectral lines in the synthetic modeling. A mock data set is used to show that absorption lines can be recovered when they are absent in one of the domains. This method can be applied to other fields that use large data sets and are currently limited by modeling accuracy. The code used in this study is made publicly available on GitHub (https://github.com/teaghan/Cycle_SN).