Integrative Sparse K-Means With Overlapping Group Lasso in Genomic Applications for Disease Subtype Discovery.

Integrative Sparse K-Means With Overlapping Group Lasso in Genomic Applications for Disease Subtype Discovery.
复制标题

DOI:
10.1214/17-aoas1033
复制
发表时间:
2017-06
期刊:
The annals of applied statistics
影响因子:
--
通讯作者:
Tseng G
Tseng G
中科院分区:
其他
文献类型:
--
作者:
Huo Z;Tseng G

文献摘要

被引文献

相似文献

癌症亚型的发现是为癌症患者提供个性化药物的第一步。随着海量多水平组学数据集的积累和已建立的生物知识库的建立,组学数据的整合与丰富的现有生物学知识的融合对于破译复杂疾病背后的生物学机制至关重要。在这篇论文中,我们提出了一种综合稀疏K-Means(IS-K Means)方法,通过稀疏重叠的群体套索,在现有生物学知识的指导下发现疾病亚型。使用交替方向乘子法(ADMM)的算法将用于快速优化。通过仿真和在乳腺癌和白血病中的三个实际应用,将IS-K均值与现有方法进行比较,展示其优越的聚类精度、特征选择、检测到的分子特征的功能标注和计算效率。
Cancer subtypes discovery is the first step to deliver personalized medicine to cancer patients. With the accumulation of massive multi-level omics datasets and established biological knowledge databases, omics data integration with incorporation of rich existing biological knowledge is essential for deciphering a biological mechanism behind the complex diseases. In this manuscript, we propose an integrative sparse K-means (is-K means) approach to discover disease subtypes with the guidance of prior biological knowledge via sparse overlapping group lasso. An algorithm using an alternating direction method of multiplier (ADMM) will be applied for fast optimization. Simulation and three real applications in breast cancer and leukemia will be used to compare is-K means with existing methods and demonstrate its superior clustering accuracy, feature selection, functional annotation of detected molecular features and computing efficiency.