Unveiling elemental fingerprints: A comparative study of clustering methods for multi-element nanoparticle data

Unveiling elemental fingerprints: A comparative study of clustering methods for multi-element nanoparticle data
复制标题

揭示元素指纹:多元素纳米颗粒数据聚类方法的比较研究

DOI:
10.1016/j.scitotenv.2023.167176
复制
发表时间:
2023
影响因子:
9.8
通讯作者:
Goharian, Erfan
Goharian, Erfan
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Erfani, Mahdi;Baalousha, Mohammed;Goharian, Erfan

文献摘要

相似文献

单粒子电感耦合等离子体飞行时间质谱仪(SP-ICP-TOF-MS)生成纳米粒子的多元素组成的大数据集。然而,从这些数据集中提取有用的信息是具有挑战性的。层次聚类(HC)已成功地应用于从SP-ICP-TOF-MS获得的多元素纳米颗粒数据中提取元素指纹。然而,许多其他聚类方法可应用于分析尚未评估的SP-ICP-TOF-MS数据。本研究通过比较三种聚类方法的性能来填补这一知识空白:HC,谱聚类和t分布随机邻居嵌入加上基于密度的空间聚类应用程序与噪声(tSNE-DBSCAN)分析SP-ICP-TOF-MS数据。这些聚类技术的性能进行了评价,通过比较提取的簇的大小和每个簇内的纳米粒子的元素组成的相似性。层次聚类往往无法实现最佳的聚类解决方案SP-ICP-TOF-MS数据,因为HC是敏感的离群值的存在。光谱聚类和tSNE-DBSCAN提取了HC未识别的聚类。这是因为谱聚类,一种基于图论的方法,揭示了数据中的全局和局部结构。tSNE减少数据并将其映射到低维空间,使DBSCAN等聚类算法能够识别元素组成存在细微差异的子聚类。然而,tSNE-DBSCAN可能会导致不满意的聚类解决方案,因为调整tSNE的困惑度超参数是一项困难且耗时的任务,并且数据点之间的相对距离没有保持。虽然这三种聚类方法成功地从SP-ICP-TOF-MS数据中提取了有用的信息,但光谱聚类通过生成大量具有相似元素组成的纳米颗粒的簇而优于HC和tSNE-DBSCAN。
Single particle-inductively coupled plasma-time of flight-mass spectrometers (SP-ICP-TOF-MS) generates large datasets of the multi-elemental composition of nanoparticles. However, extracting useful information from such datasets is challenging. Hierarchical clustering (HC) has been successfully applied to extract elemental fingerprints from multi-element nanoparticle data obtained by SP-ICP-TOF-MS. However, many other clustering approaches can be applied to analyze SP-ICP-TOF-MS data that have not yet been evaluated. This study fills this knowledge gap by comparing the performance of three clustering approaches: HC, spectral clustering, and t-distributed Stochastic Neighbor Embedding coupled with Density-Based Spatial Clustering of Applications with Noise (tSNE-DBSCAN) for analyzing SP-ICP-TOF-MS data. The performance of these clustering techniques was evaluated by comparing the size of the extracted clusters and the similarity of the elemental composition of nanoparticles within each cluster. Hierarchical clustering often failed to achieve an optimal clustering solution for SP-ICP-TOF-MS data because HC is sensitive to the presence of outliers. Spectral clustering and tSNE-DBSCAN extracted clusters that were not identified by HC. This is because spectral clustering, a method developed based on graph theory, reveals the global and local structure in the data. tSNE reduces and maps the data into a lower-dimensional space, enabling clustering algorithms such as DBSCAN to identify subclusters with subtle differences in their elemental composition. However, tSNE-DBSCAN can lead to unsatisfactory clustering solutions because tuning the perplexity hyperparameter of tSNE is a difficult and a time-consuming task, and the relative distance between datapoints is not maintained. Although the three clustering approaches successfully extract useful information from SP-ICP-TOF-MS data, spectral clustering outperforms HC and tSNE-DBSCAN by generating clusters of a large number of nanoparticles with similar elemental compositions.