Are clusterings of multiple data views independent?

Are clusterings of multiple data views independent?
复制标题

DOI:
10.1093/biostatistics/kxz001
复制
发表时间:
2020-10-01
期刊:
影响因子:
2.1
通讯作者:
Witten, Daniela
Witten, Daniela
中科院分区:
数学2区
文献类型:
--
作者:
Gao, Lucy L.;Bien, Jacob;Witten, Daniela

文献摘要

被引文献

相似文献

在先锋100 (P100)健康项目中,在多个时间点收集一组健康参与者的多种类型数据,以表征和优化健康。一种方法是在参与者中确定集群或子组,然后为每个子组量身定制个性化的健康建议。为了充分利用可用信息,很容易使用所有数据类型和时间点对参与者进行聚类。但是,基于多个数据视图对参与者进行聚类隐含地假设在所有数据视图之间共享参与者的单个底层聚类。如果这个假设不成立,那么使用多个数据视图对参与者进行聚类可能会导致错误的结果。在本文中,我们试图通过提出以下问题来评估来自不同数据视图的聚类之间存在某种潜在关系的假设:每个数据视图中的聚类是依赖的还是独立的?我们开发了一种新的测试来回答这个问题,然后我们将其应用于临床,蛋白质组学和代谢组学数据,跨越两个不同的时间点,来自P100研究。我们发现,虽然根据任何单一数据类型定义的参与者的子组似乎是随时间而变化的,但基于一种数据类型(例如蛋白质组学数据)的参与者之间的聚类似乎与基于另一种数据类型(例如临床数据)的聚类无关。
In the Pioneer 100 (P100) Wellness Project, multiple types of data are collected on a single set of healthy participants at multiple timepoints in order to characterize and optimize wellness. One way to do this is to identify clusters, or subgroups, among the participants, and then to tailor personalized health recommendations to each subgroup. It is tempting to cluster the participants using all of the data types and timepoints, in order to fully exploit the available information. However, clustering the participants based on multiple data views implicitly assumes that a single underlying clustering of the participants is shared across all data views. If this assumption does not hold, then clustering the participants using multiple data views may lead to spurious results. In this article, we seek to evaluate the assumption that there is some underlying relationship among the clusterings from the different data views, by asking the question: are the clusters within each data view dependent or independent? We develop a new test for answering this question, which we then apply to clinical, proteomic, and metabolomic data, across two distinct timepoints, from the P100 study. We find that while the subgroups of the participants defined with respect to any single data type seem to be dependent across time, the clustering among the participants based on one data type (e.g. proteomic data) appears not to be associated with the clustering based on another data type (e.g. clinical data).