Automated Identification of Characteristic Droplet Size Distributions in Stratocumulus Clouds Utilizing a Data Clustering Algorithm

Automated Identification of Characteristic Droplet Size Distributions in Stratocumulus Clouds Utilizing a Data Clustering Algorithm
复制标题

利用数据聚类算法自动识别层积云中的特征液滴尺寸分布

DOI:
10.1175/aies-d-22-0003.1
复制
发表时间:
2022
期刊:
Artificial Intelligence for the Earth Systems
影响因子:
--
通讯作者:
Shaw, Raymond A.
Shaw, Raymond A.
中科院分区:
--
文献类型:
--
作者:
Allwayin, Nithin;Larsen, Michael L.;Shaw, Alexander G.;Shaw, Raymond A.

文献摘要

相似文献

云中液滴水平的相互作用通常由一个修正的伽马函数来参数化,该伽马函数与“全球”液滴尺寸分布相适应。与微物理过程相关的“局部”液滴尺寸分布看起来像这些平均分布吗?本文描述了一种搜索和分类云中特征尺寸分布的算法。该方法结合了假设检验,特别是Kolmogorov-Smirnov(KS)检验,以及一类广泛使用的机器学习算法,用于识别具有相似属性的样本聚类:带有噪声的应用程序的基于密度的空间聚类(DBSCAN)作为具体的例子进行说明。两样本KS检验不假定任何特定的分布,是无参数的,并避免了因入库而产生的偏差。重要的是,聚类的数量不是DBSCAN类型算法的输入参数,而是以无监督的方式独立确定的。在实现时,它在来自KS测试结果的抽象空间上工作,因此集群不需要空间相关性。该方法是使用从部署在东北大西洋东部气溶胶和云实验(ACE-ENA)现场活动中的全息云探测器(HOLODEC)获得的数据进行探索的。该算法确定了存在具有几乎相同的局部大小分布的集群的证据。结果发现,云段的特征尺寸分布少则1个,多则7个。为了验证算法的稳健性,在合成数据集上进行了测试,并在合理的噪声水平下成功地识别了预定义的分布。该算法是通用的,并有望在其他应用中使用,例如对云和雨特性的遥感。意义陈述典型的云可以有数十亿滴散布在几十或数百公里的太空中。跟踪所有这些液滴的大小、位置和相互作用是不切实际的,因此,关于大小液滴的相对丰度的信息通常是用“大小分布”来量化的。然而,云中的液滴在局部相互作用,因此这项工作的动机是这样一个问题,即云滴的尺寸分布在云的不同部分是否不同。一种新的方法,基于假设检验和机器学习,确定在给定的云中包含多少不同的大小分布。这一点很重要,因为大小分布描述了云滴生长和光通过云的传输等过程。
Droplet-level interactions in clouds are often parameterized by a modified gamma fitted to a “global” droplet size distribution. Do “local” droplet size distributions of relevance to microphysical processes look like these average distributions? This paper describes an algorithm to search and classify characteristic size distributions within a cloud. The approach combines hypothesis testing, specifically, the Kolmogorov–Smirnov (KS) test, and a widely used class of machine learning algorithms for identifying clusters of samples with similar properties: density-based spatial clustering of applications with noise (DBSCAN) is used as the specific example for illustration. The two-sample KS test does not presume any specific distribution, is parameter free, and avoids biases from binning. Importantly, the number of clusters is not an input parameter of the DBSCAN-type algorithms but is independently determined in an unsupervised fashion. As implemented, it works on an abstract space from the KS test results, and hence spatial correlation is not required for a cluster. The method is explored using data obtained from the Holographic Detector for Clouds (HOLODEC) deployed during the Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA) field campaign. The algorithm identifies evidence of the existence of clusters of nearly identical local size distributions. It is found that cloud segments have as few as one and as many as seven characteristic size distributions. To validate the algorithm’s robustness, it is tested on a synthetic dataset and successfully identifies the predefined distributions at plausible noise levels. The algorithm is general and is expected to be useful in other applications, such as remote sensing of cloud and rain properties.Significance StatementA typical cloud can have billions of drops spread over tens or hundreds of kilometers in space. Keeping track of the sizes, positions, and interactions of all of these droplets is impractical, and, as such, information about the relative abundance of large and small drops is typically quantified with a “size distribution.” Droplets in a cloud interact locally, however, so this work is motivated by the question of whether the cloud droplet size distribution is different in different parts of a cloud. A new method, based on hypothesis testing and machine learning, determines how many different size distributions are contained in a given cloud. This is important because the size distribution describes processes such as cloud droplet growth and light transmission through clouds.