课题基金 / 基金详情

Semiparametric Statistical (Machine) Learning

Semiparametric Statistical (Machine) Learning
半参数统计(机器)学习
批准号:
RGPIN-2018-04868
负责人:
Loughin, Thomas
金额:
$1.68万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Loughin, Thomas的其他基金

相似基金

相关文献

中文摘要
翻译
背景*统计学习算法(SLA)是在世界范围内被用作预测复杂关系结果的现代方法的过程。SLA从一组数据中学习,该数据由一个响应(预测目标)和一个或多个解释变量(输入)组成。学习包括最小化损失-整个数据集上的预测和观察到的反应之间的距离。SLA具有称为调优参数的功能,用户可以改变这些功能以获得更合适的结果。然而,找到调整参数的最佳值的过程很繁琐,这依赖于称为交叉验证的计算密集型过程,其中整个学习过程重复多次。此外,不同的SLA可以在给定的数据集上提供略有不同的结果,并且事先不清楚应该使用哪一个。*要做的工作*拟议的研究将产生简化SLA调整过程的方法。我们还将提供将SLA组合为一个集成预测器的新方法,以利用每个单独SLA的不同优势。为此,我们将开发一种新方法,用于从数学上推导每个SLA的信息标准(IC)。ICS衡量统计模型与数据集的匹配程度,在减少损失和使模型过于复杂之间取得平衡。由于SLA不是基于统计模型,因此我们将用统计模型从数学上近似SLA(或SLA的功能),并根据该模型开发IC。模型复杂性部分很难准确估计,但我们团队在相关问题上已经成功地做到了这一点,从而产生了比现有方法更好的新的统计分析方法。*有数百种不同的SLA可以成为这项工作的目标。我们将对这些候选者中的许多人进行初步测试,以选择那些具有最佳预测质量和可在结构上服从的特征的组合来创建IC。然后,可以在单个SLA的不同版本或多个SLA上计算这些IC,以帮助选择最适合的版本。IC还可以用于模型平均,其中使用由其IC确定的权重将不同的SLA预测平均为单个预测。*结果和好处*这项研究的结果将是新的统计方法,将导致更快和更好的预测算法。这些算法将被开发成可供全国和全球数百万用户免费访问的软件。需要快速、准确地回答重要问题的用户--例如预测消费者趋势、比较潜在患者对不同疗法的反应以及预测公共政策的影响--将拥有更好的工具来执行这些任务。
英文摘要
Background***Statistical learning algorithms (SLAs) are processes that are used worldwide as modern methods for predicting results from complex relationships. SLAs learn from a set of data consisting of a response (the target for prediction) and one or more explanatory variables (the inputs). The learning consists of minimizing the loss—a distance between the predictions and the observed responses across the entire data set. SLAs have features called tuning parameters that users can vary to get a better-fitting result. However, the process of finding the best values for tuning parameters is cumbersome, relying on a computationally intensive process called cross-validation where the entire learning process is repeated numerous times. Furthermore, different SLAs can provide somewhat different results on a given data set, and it is unclear in advance which one should be used. ******Work to be done***The proposed research will generate ways to simplify the process of tuning SLAs. We will also provide new ways to combine SLAs into one ensemble predictor that makes use of the different strengths of each individual SLA. To do this, we will develop a new method for mathematically deriving an information criterion (IC) on each SLA. ICs measure the fit of a statistical model to a data set, balancing between making the loss smaller and making a model too complex. Because SLAs are not based on statistical models, we will mathematically approximate the SLA (or features of the SLA) with a statistical model, from which the IC will be developed. The model complexity component is difficult to estimate exactly, but our group has had success doing this on related problems, resulting in new statistical analysis methods with better properties than existing ones.******There are hundreds of different SLAs that could be targets for this work. We will perform preliminary tests of many of these candidates to select the ones that have the best combinations of prediction quality and structurally amenable features upon which ICs can be created. These ICs can then be computed on different versions of a single SLA or on multiple SLAs to help select the best-fitting ones. The ICs can also be for model averaging, where different SLA predictions are averaged into a single prediction using weighting determined by their ICs. ******Outcomes and Benefits ***The results of this research will be new statistical methods that lead to faster and better prediction algorithms. These algorithms will be developed into software that is freely accessible to millions of users nationwide and worldwide. Users who need fast, accurate answers to important questions—for example, forecasting consumer trends, comparing potential patient responses to different therapies, and anticipating impacts of public policies—will have better tools for performing these tasks.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Semiparametric Statistical (Machine) Learning
  • 批准号:
    RGPIN-2018-04868
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.35万
  • 财政年份:
    2022
  • 负责人:
    Loughin, Thomas
  • 依托单位:
Semiparametric Statistical (Machine) Learning
  • 批准号:
    RGPIN-2018-04868
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2021
  • 负责人:
    Loughin, Thomas
  • 依托单位:
Semiparametric Statistical (Machine) Learning
  • 批准号:
    RGPIN-2018-04868
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2020
  • 负责人:
    Loughin, Thomas
  • 依托单位:
Semiparametric Statistical (Machine) Learning
  • 批准号:
    RGPIN-2018-04868
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2019
  • 负责人:
    Loughin, Thomas
  • 依托单位:
海外基金