Robust Machine Learning Applied to Astronomical Data Sets. I. Star-Galaxy Classification of the Sloan Digital Sky Survey DR3 Using Decision Trees

Robust Machine Learning Applied to Astronomical Data Sets. I. Star-Galaxy Classification of the Sloan Digital Sky Survey DR3 Using Decision Trees
复制标题

DOI:
10.1086/507440
复制
发表时间:
2006-06
期刊:
The Astrophysical Journal
影响因子:
--
通讯作者:
N. Ball;R. Brunner;A. Myers;D. Tcheng
N. Ball;R. Brunner;A. Myers;D. Tcheng
中科院分区:
其他
文献类型:
--
作者:
N. Ball;R. Brunner;A. Myers;D. Tcheng

文献摘要

被引文献

相似文献

我们在SDSS的第三次数据发布中使用决策树对477,068个具有SDSS光谱数据的对象进行了训练,为所有1.43亿个非重复光度对象提供了分类。我们证明,这些星星/星系分类预计是可靠的约2200万个对象与r = 20。通用的机器学习环境Data-to-Knowledge和超级计算资源使得能够对决策树参数空间进行广泛的研究。这项工作提出了第一个公开发布的对象,以这种方式分类的整个SDSS数据发布。这些天体被分为星系、星星或nsng(既不是星星也不是星系),每种类别都有相应的概率。为了演示如何有效地利用这些分类,我们执行几个重要的测试。首先,我们详细的选择标准定义的概率空间内的三个类提取样本的恒星和星系到一个给定的完整性和效率。其次,我们调查的分类和外推的效果从光谱制度进行盲测试的对象在SDSS,2dFGRS,和2 QZ调查的功效。鉴于我们的光谱训练数据的光度限制,我们有效地开始推断过去r ~ 18的恒星-星系训练集。然而,通过比较我们的训练样本与分类源的数量,我们发现我们的效率似乎对r ~ 20保持稳健。因此,我们预计我们的分类对90万个星系和670万颗恒星是准确的,并且通过外推对总共800万个星系和1390万颗恒星保持稳健。
We provide classifications for all 143 million nonrepeat photometric objects in the Third Data Release of the SDSS using decision trees trained on 477,068 objects with SDSS spectroscopic data. We demonstrate that these star/galaxy classifications are expected to be reliable for approximately 22 million objects with r ≲ 20. The general machine learning environment Data-to-Knowledge and supercomputing resources enabled extensive investigation of the decision tree parameter space. This work presents the first public release of objects classified in this way for an entire SDSS data release. The objects are classified as either galaxy, star, or nsng (neither star nor galaxy), with an associated probability for each class. To demonstrate how to effectively make use of these classifications, we perform several important tests. First, we detail selection criteria within the probability space defined by the three classes to extract samples of stars and galaxies to a given completeness and efficiency. Second, we investigate the efficacy of the classifications and the effect of extrapolating from the spectroscopic regime by performing blind tests on objects in the SDSS, 2dFGRS, and 2QZ surveys. Given the photometric limits of our spectroscopic training data, we effectively begin to extrapolate past our star-galaxy training set at r ~ 18. By comparing the number counts of our training sample with the classified sources, however, we find that our efficiencies appear to remain robust to r ~ 20. As a result, we expect our classifications to be accurate for 900,000 galaxies and 6.7 million stars and remain robust via extrapolation for a total of 8.0 million galaxies and 13.9 million stars.