Detection of unknown galaxy types in large databases of galaxy images

Detection of unknown galaxy types in large databases of galaxy images
复制标题

DOI:
10.29007/5xhn
复制
发表时间:
2021-03
期刊:
--
影响因子:
--
通讯作者:
Venkat Margapuri;Basant Thapa;L. Shamir
Venkat Margapuri;Basant Thapa;L. Shamir
中科院分区:
其他
文献类型:
--
作者:
Venkat Margapuri;Basant Thapa;L. Shamir

文献摘要

相似文献

现代数字巡天利用机器人望远镜收集极其庞大的多 PB 天文数据库。虽然这些数据库可以包含数十亿个星系,但大多数星系都是已知星系类型的“常规”星系。然而,一小部分星系是罕见的、尚不为人所知的“奇特”星系。这些未知的星系具有极其重要的科学意义,但由于天文数据库规模巨大,如果没有自动化,几乎不可能找到它们。由于根据定义,这些新奇的星系是未知的,因此无法训练机器学习模型来检测它们。本文提出了一种无监督机器学习方法,用于自动检测大型数据库中的新奇星系。该方法基于一组按熵加权的大型且全面的数字图像内容描述符,并且对最远的邻居进行排序,以处理在非常大的数据集中预期的自相似的特殊星系。使用全景巡天望远镜和快速响应系统 (Pan-STARRS) 数据的实验结果表明,该方法检测新奇星系的能力优于其他浅层学习方法,如一类 SVM、局部离群因子和 K-Means,以及较新的基于深度学习的方法,如自动编码器。用于评估该方法的数据集是公开的,可以用作测试未来自动检测特殊星系算法的基准。
Modern digital sky surveys utilize robotic telescopes that collect extremely large multi-PB astronomical databases. While these databases can contain billions of galaxies, most of the galaxies are “regular” galaxies of known galaxy types. However, a small portion of the galaxies is rare “peculiar” galaxies that are not yet known. These unknown galaxies are of paramount scientific interest, but due to the enormous size of astronomical databases they are practically impossible to find without automation. Since these novelty galaxies are, by definition, not known, machine learning models cannot be trained to detect them. In this paper, an unsupervised machine learning method for automatic detection of novelty galaxies in large databases is proposed. The method is based on a large and comprehensive set of numerical image content descriptors weighted by their entropy, and the farthest neighbors are ranked-ordered to handle self-similar peculiar galaxies that are expected in the very large datasets. Experimental results using data from the Panoramic Survey Telescope and Rapid Response System (Pan-STARRS) show that the ability of the method to detect novelty galaxies outperforms other shallow learning methods such as one-class SVM, Local Outlier Factor, and K-Means, and also newer deep learning-based methods such as auto-encoders. The dataset used to evaluate the method is publicly available and can be used as a benchmark to test future algorithms for automatic detection of peculiar galaxies.