glmgraph: an R package for variable selection and predictive modeling of structured genomic data

glmgraph: an R package for variable selection and predictive modeling of structured genomic data
复制标题

DOI:
10.1093/bioinformatics/btv497
复制
发表时间:
2015-12-15
期刊:
影响因子:
5.8
通讯作者:
Chen, Jun
Chen, Jun
中科院分区:
生物学3区
文献类型:
--
作者:
Chen, Li;Liu, Han;Chen, Jun

文献摘要

被引文献

相似文献

现代高通量基因组数据分析的一个中心主题是识别相关的基因组特征,并根据个性化医疗等各种任务的选定特征建立预测模型。由于样本量 (n) 小且维度 (p) 高,将大量“组学”特征与特定表型关联起来特别具有挑战性。为了解决这个小 n、大 p 的问题,人们通过利用稀疏性假设提出了各种形式的稀疏回归模型。其中,网络约束稀疏回归模型由于其能够利用组学数据中的先验图/网络结构而受到特别关注。尽管它对于组学数据分析具有潜在的用途,但尚未公开有效的 R 实现。在这里,我们提出了一个 R 软件包“glmgraph”,它实现了稀疏线性回归和稀疏逻辑回归的图约束正则化。我们实现了用于变量选择的 L-1 罚分和极小最大凹罚分以及用于系数平滑的拉普拉斯罚分。采用高效的坐标下降算法来解决优化问题。我们通过将其应用于人类微生物组数据集来演示该包的使用,其中细菌类群的系统发育结构可用。
One central theme of modern high-throughput genomic data analysis is to identify relevant genomic features as well as build up a predictive model based on selected features for various tasks such as personalized medicine. Correlating the large number of 'omics' features with a certain phenotype is particularly challenging due to small sample size (n) and high dimensionality (p). To address this small n, large p problem, various forms of sparse regression models have been proposed by exploiting the sparsity assumption. Among these, network-constrained sparse regression model is of particular interest due to its ability to utilize the prior graph/network structure in the omics data. Despite its potential usefulness for omics data analysis, no efficient R implementation is publicly available. Here we present an R software package 'glmgraph' that implements the graph-constrained regularization for both sparse linear regression and sparse logistic regression. We implement both the L-1 penalty and minimax concave penalty for variable selection and Laplacian penalty for coefficient smoothing. Efficient coordinate descent algorithm is used to solve the optimization problem. We demonstrate the use of the package by applying it to a human microbiome dataset, where phylogeny structure among bacterial taxa is available.