Latent Model-Based Clustering for Biological Discovery

Latent Model-Based Clustering for Biological Discovery
复制标题

用于生物发现的基于潜在模型的聚类

DOI:
10.1016/j.isci.2019.03.018
复制
发表时间:
2019
期刊:
影响因子:
5.8
通讯作者:
Das, J.
Das, J.
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Bing, X;Bunea, F;Royer, M;Das, J.

文献摘要

参考文献

被引文献

相似文献

LOVE是一种强大的、可扩展的基于潜在模型的生物发现聚类方法,可以在一系列数据集上使用,以生成重叠和非重叠的聚类。在我们的公式中,集群包括与相同的潜在因素相关联的变量,并从索引我们的潜在模型的分配矩阵中确定。我们证明了分配矩阵和相应的集群是唯一定义的。我们将LOVE应用于生物数据集(基因表达,从HIV控制者和慢性进展者测量的血清学反应,疫苗诱导的体液免疫反应),从而产生有意义的生物输出。对于所有三个数据集,LOVE生成的聚类在调整参数之间保持稳定。最后,我们使用先前建立的基准将LOVE的性能与13种最先进的方法进行了比较,发现LOVE在数据集上优于这些方法。我们的研究结果表明,LOVE可以广泛用于大规模生物数据集,以生成准确和有意义的重叠和非重叠聚类。
LOVE, a robust, scalable latent model-based clustering method for biological discovery, can be used across a range of datasets to generate both overlapping and non-overlapping clusters. In our formulation, a cluster comprises variables associated with the same latent factor and is determined from an allocation matrix that indexes our latent model. We prove that the allocation matrix and corresponding clusters are uniquely defined. We apply LOVE to biological datasets (gene expression, serological responses measured from HIV controllers and chronic progressors, vaccine-induced humoral immune responses) resulting in meaningful biological output. For all three datasets, the clusters generated by LOVE remain stable across tuning parameters. Finally, we compared LOVE's performance to that of 13 state-of-the-art methods using previously established benchmarks and found that LOVE outperformed these methods across datasets. Our results demonstrate that LOVE can be broadly used across large-scale biological datasets to generate accurate and meaningful overlapping and non-overlapping clusters.
DOI: 10.1214/19-aos1877
发表时间: 2020-08-01
影响因子: 4.5
作者:
Bing, Xin;Bunea, Florentina;Wegkamp, Marten
通讯作者: Wegkamp, Marten
DOI: 10.1093/nar/gkt439
发表时间: 2013-07
影响因子: 14.9
作者:
Wang J;Duncan D;Shi Z;Zhang B
通讯作者: Zhang B