An integrative framework for Bayesian variable selection with informative priors for identifying genes and pathways.

An integrative framework for Bayesian variable selection with informative priors for identifying genes and pathways.
复制标题

DOI:
10.1371/journal.pone.0067672
复制
发表时间:
2013
期刊:
影响因子:
3.7
通讯作者:
Yang X
Yang X
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Peng B;Zhu D;Ander BP;Zhang X;Xue F;Sharp FR;Yang X

文献摘要

参考文献

被引文献

相似文献

遗传或基因组标记的发现在个性化医疗的发展中起着核心作用。在处理高维数据集时存在一个显著的挑战,因为在相对少量的受试者中收集了数千个基因或数百万个遗传变异。传统的基因明智的选择方法,使用单变量分析面临着困难,将相关性,结构或功能结构之间的分子措施。对于微阵列基因表达数据,我们首先总结了解决方案,在处理“大p,小n”的问题,然后提出了一个综合贝叶斯变量选择(iBVS)的框架,同时确定因果或标记基因和调控途径。一种新的偏最小二乘(PLS)g-先验的iBVS的开发,允许基因-基因相互作用或功能关系的先验知识。从系统生物学的角度来看,iBVS使用户能够在分层建模图中直接靶向多个基因和途径的联合作用,以预测疾病状态或表型。估计的后验选择概率提供了概率和生物学的解释。模拟数据和一组预测中风状态的微阵列数据都用于验证iBVS在具有二元结果的Probit模型中的性能。iBVS通过结合基于数据的统计和基于知识的先验知识,为有效发现各种分子生物标志物提供了一个通用框架。还讨论了关于后验推断,确定贝叶斯显著性水平和提高计算效率的指导方针。
The discovery of genetic or genomic markers plays a central role in the development of personalized medicine. A notable challenge exists when dealing with the high dimensionality of the data sets, as thousands of genes or millions of genetic variants are collected on a relatively small number of subjects. Traditional gene-wise selection methods using univariate analyses face difficulty to incorporate correlational, structural, or functional structures amongst the molecular measures. For microarray gene expression data, we first summarize solutions in dealing with ‘large p, small n’ problems, and then propose an integrative Bayesian variable selection (iBVS) framework for simultaneously identifying causal or marker genes and regulatory pathways. A novel partial least squares (PLS) g-prior for iBVS is developed to allow the incorporation of prior knowledge on gene-gene interactions or functional relationships. From the point view of systems biology, iBVS enables user to directly target the joint effects of multiple genes and pathways in a hierarchical modeling diagram to predict disease status or phenotype. The estimated posterior selection probabilities offer probabilitic and biological interpretations. Both simulated data and a set of microarray data in predicting stroke status are used in validating the performance of iBVS in a Probit model with binary outcomes. iBVS offers a general framework for effective discovery of various molecular biomarkers by combining data-based statistics and knowledge-based priors. Guidelines on making posterior inferences, determining Bayesian significance levels, and improving computational efficiencies are also discussed.
DOI: 10.1186/1471-2105-10-72
发表时间: 2009-02-26
期刊: BMC bioinformatics
影响因子: 3
作者:
Annest A;Bumgarner RE;Raftery AE;Yeung KY
通讯作者: Yeung KY
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1007/s11749-006-0040-8
发表时间: 2008-11-01
期刊: TEST
影响因子: 1.3
作者:
Cano, J. A.;Salmeron, D.;Robert, C. P.
通讯作者: Robert, C. P.
DOI: 10.1111/j.1467-9876.2005.05593.x
发表时间: 2005-01-01
影响因子: 1.6
作者:
Do, KA;Müller, P;Tang, F
通讯作者: Tang, F
DOI: 10.1093/bioinformatics/17.6.509
发表时间: 2001-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Baldi, P;Long, AD
通讯作者: Long, AD