Ensemble Methods for Classification/Prediction With High-Dimensional Explanatory Variables
Ensemble Methods for Classification/Prediction With High-Dimensional Explanatory Variables
批准号:
RGPIN-2014-04962
负责人:
Welch, William
金额:
$1.31万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Advances in science and engineering have vastly increased the number of variables available to predict / classify a response outcome of interest. At the same time the information in the data may be sparse. Novel methods based on ensembles of models are proposed for higher prediction accuracy. Methodology will be developed for two problems with these characteristics: prediction of complex computer codes and prediction / classification in analysis of drug discovery data.**Deterministic computer models can have complex relationships with high-dimensional input (explanatory) variables. For instance, the Community Land Model of the carbon cycle and vegetation dynamics has hundreds of inputs for the ecosystem, climate, hydrology, etc. Experiments with about 100 variables are aimed at sensitivity analysis, i.e., find the inputs that have most impact on an output such as a measure of total vegetation. It is feasible to make thousands of computer model runs, yet work to date shows the input-output relationships are hard to model with useful accuracy. Most likely, there are complex interaction effects between the inputs, and identifying them is a challenge because of the high dimensionality.**In drug discovery, the input variables are "chemical descriptors" from computational chemistry to characterize drug-like molecules. Many sets are available, and each can have thousands of variables. The response variable or output is from a physical assay of activity against a biological target implicated in a disease. A statistical model relating biological activity to the chemical inputs can be used to predict activity of molecules that have not been assayed yet, increasing efficiency of the process to search for candidate drugs. Unfortunately, active molecules are rare, so there is a paucity of information in the response data to fit a model.**Gaussian Processes (GPs) are widely used to model the deterministic input-output relationship of a computer code. They have also been used in the analysis of drug discovery data. The proposed approach to high-dimensional input and limited data information is based on ensembles of GPs, either by building separate models and averaging them, or by ensembles of correlation functions (which are key to the GP approach). Ensembles have well known general advantages in prediction accuracy and are established as among the best for the drug discovery problem, for example. They typically generate multiple prediction models by perturbing the data (bootstrapping) or dividing the data observations and then fitting a model to each data set created. The models are then averaged when making predictions. With high-dimensional input, however, sparse information in the response data means that most of the input variables are unused in a model when it is fit to data. **In contrast, the proposed approach is to build an ensemble of models over distinct subsets of input variables. A subset of inputs with interaction effects should be in the same model; variables that do not interact can be in separate models. It is easier to fill the input space in a data set densely a few variables at a time, increasing prediction accuracy. Furthermore, by attributing variables to different models, more inputs have a chance to contribute to prediction accuracy. The challenges and goals of the research program are how to identify subsets of high-dimensional input variables that should be together in the same model, how to combine models for high overall prediction accuracy, and efficient algorithms to overcome the computational demands of GP models. The over-arching goal is to understand how a statistical model like a GP should be tuned to the complexities of relationships involving high-dimensional input.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Adaptive Design for Fast Machine/Statistical Learning
-
批准号:RGPIN-2019-05019
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2022
-
负责人:Welch, William
-
依托单位:
Adaptive Design for Fast Machine/Statistical Learning
-
批准号:RGPIN-2019-05019
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2021
-
负责人:Welch, William
-
依托单位:
Adaptive Design for Fast Machine/Statistical Learning
-
批准号:RGPIN-2019-05019
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2020
-
负责人:Welch, William
-
依托单位:
Adaptive Design for Fast Machine/Statistical Learning
-
批准号:RGPIN-2019-05019
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2019
-
负责人:Welch, William
-
依托单位:
Ensemble Methods for Classification/Prediction With High-Dimensional Explanatory Variables
-
批准号:RGPIN-2014-04962
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2017
-
负责人:Welch, William
-
依托单位:
Ensemble Methods for Classification/Prediction With High-Dimensional Explanatory Variables
-
批准号:RGPIN-2014-04962
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2016
-
负责人:Welch, William
-
依托单位:
Ensemble Methods for Classification/Prediction With High-Dimensional Explanatory Variables
-
批准号:RGPIN-2014-04962
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2015
-
负责人:Welch, William
-
依托单位:
Ensemble Methods for Classification/Prediction With High-Dimensional Explanatory Variables
-
批准号:RGPIN-2014-04962
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2014
-
负责人:Welch, William
-
依托单位:
Classification: methodology for variable selection and efficient tuning and comparasion of models
-
批准号:36462-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2012
-
负责人:Welch, William
-
依托单位:
Classification: methodology for variable selection and efficient tuning and comparasion of models
-
批准号:36462-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2011
-
负责人:Welch, William
-
依托单位:
Classification: methodology for variable selection and efficient tuning and comparasion of models
-
批准号:36462-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2010
-
负责人:Welch, William
-
依托单位:
Classification: methodology for variable selection and efficient tuning and comparasion of models
-
批准号:36462-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2009
-
负责人:Welch, William
-
依托单位:
Classification: methodology for variable selection and efficient tuning and comparasion of models
-
批准号:36462-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2008
-
负责人:Welch, William
-
依托单位:
Bayesian analysis of computer experiments
-
批准号:36462-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.97万
-
财政年份:2007
-
负责人:Welch, William
-
依托单位:
Bayesian analysis of computer experiments
-
批准号:36462-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.97万
-
财政年份:2006
-
负责人:Welch, William
-
依托单位:
Bayesian analysis of computer experiments
-
批准号:36462-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.97万
-
财政年份:2005
-
负责人:Welch, William
-
依托单位:
Bayesian analysis of computer experiments
-
批准号:36462-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.97万
-
财政年份:2004
-
负责人:Welch, William
-
依托单位:
Bayesian analysis of computer experiments
-
批准号:36462-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.97万
-
财政年份:2003
-
负责人:Welch, William
-
依托单位:
Strategies for Collection and Analysis of High Throughput Screening Data in Drug Discovery
-
批准号:246312-2001
-
项目类别:Strategic Projects - Group
-
资助金额:$3.28万
-
财政年份:2003
-
负责人:Welch, William
-
依托单位:
Methodology for computer experiments
-
批准号:36462-1999
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.53万
-
财政年份:2002
-
负责人:Welch, William
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: