Clustering Sparse Data With Feature Correlation With Application to Discover Subtypes in Cancer.

Clustering Sparse Data With Feature Correlation With Application to Discover Subtypes in Cancer.
复制标题

将稀疏数据与特征相关性进行聚类以发现癌症亚型

DOI:
10.1109/access.2020.2982569
复制
发表时间:
2020
期刊:
影响因子:
3.9
通讯作者:
Chen, Ping
Chen, Ping
中科院分区:
计算机科学3区
文献类型:
--
作者:
Qiang, Jipeng;Ding, Wei;Kuijjer, Marieke;Quackenbush, John;Chen, Ping

文献摘要

参考文献

被引文献

相似文献

本文针对具有高维特征的数据,研究了如何通过考虑特征交互网络来计算两个样本之间的相似度问题,其中特征交互网络表示特征之间的关系。这与一些传统的方法不同,传统的方法是基于代表样本之间关系的样本网络来学习相似性。因此,我们提出了一种新的基于网络的相似性度量来计算样本之间的相似性,该度量融合了特征交互网络的知识,以克服数据稀疏性问题。我们的相似性度量使用了一种新的特征对齐相似性度量,它不直接计算样本之间的相似性,而是将每个样本投影到一个特征交互网络中,并使用网络中样本顶点之间的相似性来测量两个样本之间的相似性。因此,当两个样本没有任何共同特征时,当它们的特征共享相似的网络区域时,它们可能具有更高的相似值。为了确保该度量在实际应用中是有用的,我们通过结合基因相互作用网络的信息,将我们的度量应用于肿瘤突变数据中的亚型发现。我们使用合成数据和真实肿瘤突变数据的实验结果表明,我们的方法在癌症亚型发现方面优于顶级竞争对手。此外,我们的方法可以识别真实癌症数据中其他聚类算法无法检测到的癌症亚型。
In this paper, given data with high-dimensional features, we study this problem of how to calculate the similarity between two samples by considering feature interaction network, where a feature interaction network represents the relationship between features. This is different from some traditional methods, those of which learn similarities based on a sample network that represents the relationship between samples. Therefore, we propose a novel network-based similarity metric for computing the similarity between samples, which incorporates the knowledge of feature interaction network, in order to overcome the data sparseness problem. Our similarity metric uses a new Feature Alignment Similarity measure, which does not directly compute the similarities among samples, but projects each sample into a feature interaction network and measures the similarities between two samples using the similarities between the vertices of the samples in the network. As such, when two samples do not share any common features, they are likely to have higher similarity values when their features share the similar network regions. For ensuring that the metric is useful in a real-world application, we apply our metric to discover subtypes in tumor mutational data by incorporating the information of the gene interaction network. Our experimental results from using synthetic data and real-world tumor mutational data show that our approach outperforms the top competitors in cancer subtype discovery. Furthermore, our approach can identify cancer subtypes that cannot be detected by other clustering algorithms in real cancer data.
DOI: 10.1073/pnas.0307752101
发表时间: 2004-04-06
影响因子: 11.1
作者:
Griffiths, TL;Steyvers, M
通讯作者: Steyvers, M
DOI: 10.1186/s13015-019-0146-7
发表时间: 2019-03-30
影响因子: 1
作者:
Hajkarim, Morteza Chalabi;Upfal, Eli;Vandin, Fabio
通讯作者: Vandin, Fabio
DOI: 10.1073/pnas.0308531101
发表时间: 2004-03-23
影响因子: 11.1
作者:
Brunet, JP;Tamayo, P;Mesirov, JP
通讯作者: Mesirov, JP
DOI: 10.1101/gr.118992.110
发表时间: 2011-07-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Lee, Insuk;Blom, U. Martin;Marcotte, Edward M.
通讯作者: Marcotte, Edward M.
DOI: 10.1162/jmlr.2003.3.4-5.993
发表时间: 2003-05-15
影响因子: 6
作者:
Blei, DM;Ng, AY;Jordan, MI
通讯作者: Jordan, MI