The usefulness of sparse k-means in metabolomics data: An example from breast cancer data

The usefulness of sparse k-means in metabolomics data: An example from breast cancer data
复制标题

DOI:
10.1101/2022.02.05.479235
复制
发表时间:
2022-02
期刊:
bioRxiv
影响因子:
--
通讯作者:
Misa Goudo;M. Sugimoto;S. Hiwa;Tomoyuki Hiroyasu
Misa Goudo;M. Sugimoto;S. Hiwa;Tomoyuki Hiroyasu
中科院分区:
其他
文献类型:
--
作者:
Misa Goudo;M. Sugimoto;S. Hiwa;Tomoyuki Hiroyasu

文献摘要

相似文献

在处理代谢组学数据时,来自数千种代谢物的多维定量数据通常是稀疏的,也就是说,只有一小部分代谢物与感兴趣的表型相关。因此,聚类用于从组学数据中发现亚型。稀疏处理是一种有效的聚类技术,它从组学数据中选择重要的代谢产物。本研究探讨了稀疏k均值算法对代谢组学数据的有效性。具体而言,稀疏k均值用于聚类两项研究中乳腺癌患者的血脂代谢物数据:(1)绝经前和绝经后,以及(2)术前和术后化疗。在这两种情况下,稀疏k-均值显示出相当的判别准确性,代谢物比k-均值少。此外,当L1范数值变化时,未观察到显著变化。稀疏k-means和k-means的平均轮廓系数分别为(1)0.38 ± 0.14(S.D.)0.17 ± 0.01,(2)0.38 ± 0.07和0.17 ± 0.01,表明使用稀疏k均值进行特征选择可以改善聚类结果。此外,无论试验数据或L1范数的约束值如何,使用稀疏k均值的代谢物选择都是一致的,表明具有稳健性。
In processing metabolomics data, multidimensional quantitative data from thousands of metabolites are often sparse, that is, only a small fraction of metabolites are relevant to the phenotype of interest. Clustering is therefore used to discover subtypes from omics data. Sparse processing, which selects important metabolites from the total omics data, is an effective clustering technique. This study investigated the effectiveness of sparse k-means for metabolomics data. Specifically, sparse k-means was used to cluster blood lipid metabolite data of breast cancer patients in two studies: (1) before and after menopause, and (2) pre- and postoperative chemotherapy. In both cases, sparse k-means showed comparable discrimination accuracy with fewer metabolites than k-means. Furthermore, when the L1 norm values were varied, no significant changes were observed. The mean silhouette coefficients of sparse k-means and k-means were (1) 0.38 ± 0.14 (S.D.) and 0.17 ± 0.01, (2) 0.38 ± 0.07 and 0.17 ± 0.01, indicating that feature selection using sparse k-means can improve clustering results. In addition, metabolite selection using sparse k-means was consistent regardless of the test data or the constrained value of the L1 norm, indicating robustness.