Dimensionality Reduction of SDSS Spectra with Variational Autoencoders

Dimensionality Reduction of SDSS Spectra with Variational Autoencoders
复制标题

DOI:
10.3847/1538-3881/ab9644
复制
发表时间:
2020-02
期刊:
The Astronomical Journal
影响因子:
--
通讯作者:
S. Portillo;J. Parejko;J. Vergara;A. Connolly
S. Portillo;J. Parejko;J. Vergara;A. Connolly
中科院分区:
其他
文献类型:
--
作者:
S. Portillo;J. Parejko;J. Vergara;A. Connolly

文献摘要

被引文献

相似文献

高分辨率星系光谱包含了许多关于星系物理的信息,但这些光谱的高维性使得它们所包含的信息很难得到充分利用。我们应用变分自动编码器(VAEs),这是一种非线性降维技术,用于斯隆数字巡天(SDSS)的光谱样本。与广泛使用的主成分分析(PCA)不同,VAEs可以捕捉潜在参数与数据之间的非线性关系。我们发现,VAE仅用6个潜参数就能很好地重建SDSS谱,优于相同分量个数的主成分分析。在这个潜在的空间中,不同的星系类别被自然地分开,而没有给VAE赋予类别标签。VAE潜在空间是可解释的,因为VAE可用于在潜在空间中的任何点进行合成光谱。例如,沿着潜在空间的轨迹制作合成光谱可以产生在两个不同类型的星系之间插入的真实光谱序列。使用潜在空间来发现离群值可能会产生有趣的光谱:在我们的小样本中,我们立即发现不寻常的数据制品和被错误归类为星系的恒星。在这项探索性工作中,我们展示了VAE创建了紧凑的、可解释的潜在空间,这些空间捕捉了数据的非线性特征。虽然VAE需要相当长的时间来训练(48,000个光谱的≈1天),但一旦接受训练,VAE就可以实现对大型天文数据集的快速探索。
High-resolution galaxy spectra contain much information about galactic physics, but the high dimensionality of these spectra makes it difficult to fully utilize the information they contain. We apply variational autoencoders (VAEs), a nonlinear dimensionality reduction technique, to a sample of spectra from the Sloan Digital Sky Survey (SDSS). In contrast to principal component analysis (PCA), a widely used technique, VAEs can capture nonlinear relationships between latent parameters and the data. We find that a VAE can reconstruct the SDSS spectra well with only six latent parameters, outperforming PCA with the same number of components. Different galaxy classes are naturally separated in this latent space, without class labels having been given to the VAE. The VAE latent space is interpretable because the VAE can be used to make synthetic spectra at any point in latent space. For example, making synthetic spectra along tracks in latent space yields sequences of realistic spectra that interpolate between two different types of galaxies. Using the latent space to find outliers may yield interesting spectra: in our small sample, we immediately find unusual data artifacts and stars misclassified as galaxies. In this exploratory work, we show that VAEs create compact, interpretable latent spaces that capture nonlinear features of the data. While a VAE takes substantial time to train (≈1 day for 48,000 spectra), once trained, VAEs can enable the fast exploration of large astronomical data sets.