Cluster analysis of massive datasets in astronomy

Cluster analysis of massive datasets in astronomy
复制标题

天文学海量数据集的聚类分析

DOI:
10.1007/s11222-007-9027-x
复制
发表时间:
2007
影响因子:
2.2
通讯作者:
M. Hendry
M. Hendry
中科院分区:
数学2区
文献类型:
--
作者:
Woncheol Jang;M. Hendry

文献摘要

被引文献

相似文献

摘要 星系团是追踪宇宙中质量分布的有用代用品。通过测量不同尺度上星系团的质量,人们可以跟踪质量分布的演变(Martínez和Saar,Statistics of the Galaxy Distribution,2002)。可以证明,寻找星系团等价于寻找密度轮廓星团(Hartigan,聚类算法,1975):水平集Sc {f>c}的连通分量,其中f是概率密度函数。Cuevas等人(可以。J. Stat. 28,367-382,2000;计算。Stat.数据分析36,441-459,2001)提出了一种用于密度轮廓聚类的非参数方法,试图通过最小生成树来寻找密度轮廓聚类。虽然他们的算法在概念上很简单,但它需要对大型数据集进行密集的计算。我们提出了一个更有效的聚类方法的基础上,他们的算法与快速傅立叶变换(FFT)。该方法被应用于大型天文巡天数据的星系聚类研究。
Abstract Clusters of galaxies are a useful proxy to trace the distribution of mass in the universe. By measuring the mass of clusters of galaxies on different scales, one can follow the evolution of the mass distribution (Martínez and Saar, Statistics of the Galaxy Distribution, 2002). It can be shown that finding galaxy clusters is equivalent to finding density contour clusters (Hartigan, Clustering Algorithms, 1975): connected components of the level set Sc≡{f>c} where f is a probability density function. Cuevas et al. (Can. J. Stat. 28, 367–382, 2000; Comput. Stat. Data Anal. 36, 441–459, 2001) proposed a nonparametric method for density contour clusters, attempting to find density contour clusters by the minimal spanning tree. While their algorithm is conceptually simple, it requires intensive computations for large datasets. We propose a more efficient clustering method based on their algorithm with the Fast Fourier Transform (FFT). The method is applied to a study of galaxy clustering on large astronomical sky survey data.