A Sparse PLS for Variable Selection when Integrating Omics Data

A Sparse PLS for Variable Selection when Integrating Omics Data
复制标题

DOI:
10.2202/1544-6115.1390
复制
发表时间:
2008-01-01
影响因子:
0.9
通讯作者:
Besse, Philippe
Besse, Philippe
中科院分区:
数学4区
文献类型:
--
作者:
Le Cao, Kim-Anh;Rossouw, Debra;Besse, Philippe

文献摘要

被引文献

相似文献

最近的生物技术进步允许多种类型的组学数据,例如转录组、蛋白质组或代谢组数据集被整合。特征选择的问题在分类方面已经讨论了几次,但在整合数据时需要以特定的方式处理。在这项研究中,我们关注的是在相同样本上测量的两个区块数据的整合。我们的目标是将两个数据集的整合和同时变量选择结合在一个步骤中,使用偏最小二乘回归(偏最小二乘回归)变量变量来方便生物学家的解释。针对这些新出现的问题,引入了一种称为“稀疏偏最小二乘法”的计算方法来进行预测分析。在计算奇异值分解时,通过对偏最小二乘加载向量进行Lasso惩罚,实现了该方法的稀疏性,证明了稀疏偏最小二乘方法的有效性和生物学意义。在模拟数据集和真实数据集上与经典的最小二乘法进行了比较。在一个数据集上,对所获得的结果提供了全面的生物学解释。我们证明了稀疏偏最小二乘为高维数据集提供了一种有价值的变量选择工具。
Recent biotechnology advances allow for multiple types of omics data, such as transcriptomic, proteomic or metabolomic data sets to be integrated. The problem of feature selection has been addressed several times in the context of classification, but needs to be handled in a specific manner when integrating data. In this study, we focus on the integration of two-block data that are measured on the same samples. Our goal is to combine integration and simultaneous variable selection of the two data sets in a one-step procedure using a Partial Least Squares regression (PLS) variant to facilitate the biologists' interpretation. A novel computational methodology called "sparse PLS" is introduced for a predictive analysis to deal with these newly arisen problems. The sparsity of our approach is achieved with a Lasso penalization of the PLS loading vectors when computing the Singular Value Decomposition.Sparse PLS is shown to be effective and biologically meaningful. Comparisons with classical PLS are performed on a simulated data set and on real data sets. On one data set, a thorough biological interpretation of the obtained results is provided. We show that sparse PLS provides a valuable variable selection tool for highly dimensional data sets.