Quasar and galaxy classification in Gaia Data Release 2

Quasar and galaxy classification in Gaia Data Release 2
复制标题

DOI:
10.1093/mnras/stz2947
复制
发表时间:
2019-10
影响因子:
4.8
通讯作者:
C. Bailer-Jones;M. Fouesneau;R. Andrae
C. Bailer-Jones;M. Fouesneau;R. Andrae
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
C. Bailer-Jones;M. Fouesneau;R. Andrae

文献摘要

被引文献

相似文献

我们构建了一个基于高斯混合模型的监督分类器,仅使用该版本中的光度和天体测量数据对盖亚数据版本 2 (GDR2) 中的对象进行概率分类。该模型经过经验训练,可将物体分为三类:恒星、类星体、星系,G ≥ 14.5 星等,直至盖亚星等限制 G = 21.0 星等。通过与来自斯隆数字巡天的光谱分类的物体进行交叉匹配,为训练集识别星系和类星体。星星是直接从 GDR2 定义的。当考虑到类星体比恒星稀有 500 倍、星系比恒星稀有 7500 倍(类不平衡问题)时,以阈值概率 0.5 分类的样本预计类星体的纯度为 0.43,星系的纯度为 0.28,完整性分别为 0.58 和 0.72。通过采用更高的阈值,纯度可以提高到0.60。如果不考虑河外物体(先验类别)的预期低频率,将会给出错误乐观的性能预测和严重不纯的样本。将我们的模型应用于 GDR2 中所有具有所需特征的 12 亿个物体,我们将 230 万个物体分类为类星体,将 37 万个物体分类为星系(个体概率高于 0.5)。星系数量较少是由于卫星探测算法和地面数据选择对扩展天体存在强烈偏差。我们推断类星体和星系的真实数量(这些类别是由我们的训练集定义的)分别为 690 000 个和 110 000 个(±50%)。这项工作的目的是了解仅使用 GDR2 数据对河外物体进行分类的效果如何。通过为 GDR3 计划的低分辨率光谱 (BP/RP),应该可以实现更好的分类。
We construct a supervised classifier based on Gaussian Mixture Models to probabilistically classify objects in Gaia data release 2 (GDR2) using only photometric and astrometric data in that release. The model is trained empirically to classify objects into three classes – star, quasar, galaxy – for G ≥ 14.5 mag down to the Gaia magnitude limit of G = 21.0 mag. Galaxies and quasars are identified for the training set by a cross-match to objects with spectroscopic classifications from the Sloan Digital Sky Survey. Stars are defined directly from GDR2. When allowing for the expectation that quasars are 500 times rarer than stars, and galaxies 7500 times rarer than stars (the class imbalance problem), samples classified with a threshold probability of 0.5 are predicted to have purities of 0.43 for quasars and 0.28 for galaxies, and completenesses of 0.58 and 0.72, respectively. The purities can be increased up to 0.60 by adopting a higher threshold. Not accounting for this expected low frequency of extragalactic objects (the class prior) would give both erroneously optimistic performance predictions and severely impure samples. Applying our model to all 1.20 billion objects in GDR2 with the required features, we classify 2.3 million objects as quasars and 0.37 million objects as galaxies (with individual probabilities above 0.5). The small number of galaxies is due to the strong bias of the satellite detection algorithm and on-ground data selection against extended objects. We infer the true number of quasars and galaxies – as these classes are defined by our training set – to be 690 000 and 110 000, respectively (±50 per cent). The aim of this work is to see how well extragalactic objects can be classified using only GDR2 data. Better classifications should be possible with the low resolution spectroscopy (BP/RP) planned for GDR3.