Response to ‘Letter to the Editor: on the stability and internal consistency of component-wise sparse mixture regression based clustering’, Zhang et al.

Response to ‘Letter to the Editor: on the stability and internal consistency of component-wise sparse mixture regression based clustering’, Zhang et al.
复制标题

回应“致编辑的信:关于基于组件的稀疏混合回归聚类的稳定性和内部一致性”,Zhang 等人。

DOI:
10.1093/bib/bbac262
复制
发表时间:
2022
影响因子:
9.5
通讯作者:
Cao, Sha
Cao, Sha
中科院分区:
生物学2区
文献类型:
--
作者:
Chang, Wennan;Zhang, Chi;Cao, Sha

文献摘要

相似文献

我们最近发表了一种聚类方法,即CSMR,基于高维预测变量背景下的混合回归[1]。我们的工作的动机是提供一个可扩展的方法来处理高维分子特征时,存在异质性的分子特征和感兴趣的疾病表型之间的关系。在我们最初的工作中,我们模拟了不同的数据环境,包括样本大小N,聚类区分和聚类特定预测因子的数量M0,聚类数量K和噪声水平σ。我们在模拟数据集上评估了CSMR,基于其使用兰德指数聚类的准确性,使用真阳性/阴性率进行特征选择,以及使用Pearson相关性进行预测和观察响应值的一致性。在现实世界的癌细胞系百科全书(CCLE)数据集[2]中,我们使用交叉验证基于预测和观察到的响应值的一致性评估了CSMR。Zhang等人的信提出了这样的担忧:当簇区分预测因子M0的数量从5增加到20时,CSMR的性能显着下降,如通过包括调整的兰德指数(ARI)和聚类内部一致性(IC)的评估度量所证明的,这两者都由Zhang等人提倡用于评估聚类方法。我们感谢张等人。的利益,我们的方法,并同意他们的意见,性能下降,在大M0的情况下确实是真实的。我们想提到的是,混合回归模型,同时存在低维和高维特征,依赖于期望最大化(EM)算法或其变体的解决方案。然而,EM算法往往局部收敛于混合参数的极大似然估计[3],并且使用
We have recently published a clustering method, namely CSMR, based on mixture regression in the context of high-dimensional predictors [1]. The motivation of our work was to provide a scalable method to deal with highdimensional molecular features when there exist heterogeneous relationships between the molecular features and a disease phenotype of interest. In our original work, we simulated different data environments, including sample size N, number of cluster-differentiating and cluster-specific predictors M0, number of clusters K, and noise level σ. We evaluated CSMR on the simulation datasets, based on its accuracy of clustering using Rand index, and feature selection using true positive/negative rate, as well as the consistency of predicted and observed response values using Pearson correlation. In a realworld Cancer Cell Line Encyclopedia (CCLE) dataset [2], we evaluated CSMR based on the consistency of predicted and observed response values using crossvalidation.The letter by Zhang et al. raised the concerns that the performance of CSMR drops significantly when the number of cluster-differentiating predictors M0 increases from 5 to 20, as demonstrated by evaluation metrics including the adjusted Rand index (ARI) and the clustering internal consistency (IC), both of which are advocated by Zhang et al. to be used for evaluating clustering methods. We appreciate Zhang et al.’s interests in our method and agree that their observations of the performance drop in the large M0 case are indeed true. We would like to mention that mixture regression models, with the presence of both low-and high-dimensional features, rely on Expectation–Maximization (EM) algorithm or its variants for a solution. However, EM algorithm often converges to the maximum likelihood estimate of the mixture parameters locally [3], and using