Bayesian Variable Selection in Generalized Linear Models with Missing Varibles
Bayesian Variable Selection in Generalized Linear Models with Missing Varibles
批准号:
8471550
负责人:
XIAOWEI YANG
金额:
$23.0万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-08-11 至 2014-08-30
关键词:
AccountingAddressAlgorithmsArchivesAutistic DisorderBayesian MethodBinomial DistributionBiomedical ResearchChild AbuseClinical TrialsComplexComputer softwareDataData AnalysesData SetDependenceDevelopmentDropoutEffectivenessEquationEvaluationGeneric DrugsGenesGenetic TranscriptionImmunobiologyIndividualLibrariesLinear ModelsLinear RegressionsMarkov ChainsMedical ResearchMethodologyMethodsModelingObservational StudyOutcomePathway interactionsPerformancePhenotypePoisson DistributionProblem behaviorProceduresProcessResearchResearch PersonnelResortSchemeSideSocial ProblemsSocietiesSolutionsStatistical Data InterpretationStructureTestingUncertaintybasebehavior measurementclinically relevantcytokineempoweredflexibilityimprovedsmoking cessationsoftware developmenttool
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): In conducting medical research, especially with behavioral and social problems, a challenge for statistical data analysis comes from the problems introduced by missing values. Missing values may be caused by subjective (e.g., nonresponse and dropout) and technical reasons (e.g., censoring over/below quantization level). Generalized linear models (GLMs) are popularly applied in biomedical data analysis where a fundamental task is to interpret or predict an outcome variable by a subset of potentially explanatory variables. Given an incomplete data set, practitioners frequently resort to the strategy of case-deletion where individuals are excluded from consideration if they miss any of the variables targeted for analysis. This is the default option used in many software packages. Yet, case-deletion may not only sacrifice useful information, but also give rise to biased estimates because it requires strong assumptions on the missingness mechanisms. A more satisfactory solution for missing data problems involves multiple imputation, where several imputations are created for the same set of missing values. The variance between imputations reflects the uncertainty due to missingness. Across multiply imputed data sets, however, traditional variable selection methods (based on significance tests or various criteria) often result in models with different selected predictors, thus presenting a problem of combining the models to make final inferences. In this R01 proposal with a 3-year research plan, we aim to develop two alternative strategies of variable selection for GLMs with missing values by drawing on a Bayesian framework. One approach, which we call "impute, then select" (ITS) involves initially performing multiple imputation and then applying Bayesian variable selection to the multiply imputed data sets. The second strategy - "simultaneously impute and select" (SIAS) - is to conduct Bayesian variable selection and missing data imputation simultaneously within one Markov Chain Monte Carlo (MCMC) process. ITS and SIAS offer two generic frameworks within which various Bayesian variable selection algorithms and missing data imputation algorithms can be implemented. Both strategies will be developed, evaluated, and implemented into an R library for normal regression, binomial regression, and other GLMs with categorical and/or continuous explanatory variables. Practical data sets from several studies on substances abuse and childhood autism will be used to address the effectiveness and flexibility of the proposed strategies. Development of these procedures and contribution of the software to statisticians and researchers in medical research would significantly improve the quality of evaluation of important and clinically relevant data.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1371/journal.pone.0062129
发表时间:
2013
期刊:
PloS one
影响因子:
3.7
作者:
[Zhang X, Yang X, Yuan Z, Liu Y, Li F, Peng B, Zhu D, Zhao J, Xue F]
通讯作者:
Xue F
DOI:
10.1371/journal.pone.0067672
发表时间:
2013
期刊:
PloS one
影响因子:
3.7
作者:
[Peng B, Zhu D, Ander BP, Zhang X, Xue F, Sharp FR, Yang X]
通讯作者:
Yang X
Integrative Bayesian variable selection with gene-based informative priors for genome-wide association studies.
综合贝叶斯变量选择与基于基因的信息先验,用于全基因组关联研究。
DOI:
10.1186/s12863-014-0130-7
发表时间:
2014-12-10
期刊:
BMC genetics
影响因子:
2.9
作者:
[Zhang X, Xue F, Liu H, Zhu D, Peng B, Wiemels JL, Yang X]
通讯作者:
Yang X
Bayesian Variable Selection in Generalized Linear Models with Missing Varibles
-
批准号:8317303
-
项目类别:
-
资助金额:$9.69万
-
财政年份:2011
-
负责人:XIAOWEI YANG
-
依托单位:
Bayesian Variable Selection in Generalized Linear Models with Missing Varibles
-
批准号:8543193
-
项目类别:
-
资助金额:$9.54万
-
财政年份:2011
-
负责人:XIAOWEI YANG
-
依托单位:
Bayesian Variable Selection in Generalized Linear Models with Missing Varibles
-
批准号:8194802
-
项目类别:
-
资助金额:$19.27万
-
财政年份:2011
-
负责人:XIAOWEI YANG
-
依托单位:
iPhone-based Real-time Data Solution for Drug Abuse and Other Medical Research
-
批准号:7672825
-
项目类别:
-
资助金额:$9.99万
-
财政年份:2009
-
负责人:XIAOWEI YANG
-
依托单位:
Transition Model for Incomplete Longitudinal Binary Data
-
批准号:6676189
-
项目类别:
-
资助金额:$6.0万
-
财政年份:2003
-
负责人:XIAOWEI YANG
-
依托单位:
DEVELOPMENT OF AN AUTOMATED NEURAL SPIKE DISCRIMINATOR
-
批准号:3504570
-
项目类别:
-
资助金额:$5.0万
-
财政年份:1991
-
负责人:XIAOWEI YANG
-
依托单位:
海外基金